Developing Evaluation Functions: Getting Started¶
What is an Evaluation Function?¶
It's a cloud function which performs some computation given some user input (the response), a problem-specific source of truth (the answer), and some optional parameters (params). Evaluation functions capture and automate the role of a teacher who has to keep marking the same question countless times. The simplest example for this would be one which checks for exact equivalence - where the function signals a response is correct only if it is identical to the answer. However, more complex and exotic ones such as symbolic expression equivalence and parsing of physical units can be imagined.
Getting Setup for Development¶
These steps are for Python functions
They assume a function created from the Python boilerplate. The concepts (config, deploy
pipeline, µEd registration) are the same for every language, but the file layout and local
commands in steps 3–4 are Python-specific — for other languages follow
Other Languages and the chosen boilerplate's README.md.
- Get the code on your local machine (Using github desktop or the
gitcli)- For new functions: create a new repository from the
evaluation-function-boilerplate-pythontemplate via Use this template, choosing theLambda Feedbackorganisation as the owner. Make sure the new repository is set to public (it needs access to organisation secrets). Boilerplates for other languages also exist —evaluation-function-boilerplate-wolframandevaluation-function-boilerplate-lean; see Other Languages. - For existing functions: please make your changes on a new separate branch
- For new functions: create a new repository from the
-
If you are creating a new function, set its deployed name in the
config.jsonfile in the root directory:{ "EvaluationFunctionName": "myFunction" }The name must be unique across the organisation and is conventionally
lowerCamelCase. 3. You are now ready to start making changes. The function logic lives in theevaluation_function/package: 1.evaluation_function/evaluation.py: contains the mainevaluation_function, which is called to compare a response to an answer.[`evaluation.py` Specification](specification.md#evaluationpy){ .md-button }-
evaluation_function/preview.py: containspreview_function, which pre-processes a response for live display (e.g. rendered LaTeX) without grading it. -
evaluation_function/evaluation_test.py: where you test the logic inevaluation.py, usingpytest. -
evaluation_function/main.py: the entry point. It callslf_toolkit.create_server()and registers yourevaluation_functionandpreview_functionwith it. You rarely need to change this file. -
Documentation files:
-
docs/dev.md: edited to reflect any changes/features from a developer perspective. It is baked into the function's image and pulled into this site under the deployed functions section. -
docs/user.md: documents how a teacher uses the function when editing content on the LambdaFeedback platform. These files are displayed in the Teacher section. -
README.md: replace the boilerplate's generic title, description and Quickstart section with your function's own, and have it link todocs/dev.md,docs/user.mdand this site rather than restate them. See the README convention.
-
-
-
Changes can be tested locally by running your tests from the repository root:
Running and Testing Functions Locallypoetry install poetry run pytest -
The pipeline has two environments:
-
Staging — pushing to the
mainbranch triggers thestaging-deploy.ymlworkflow, which runs the test suite and (on success) builds and deploys the docker image to staging. -
Production — once you are happy with the staging deployment, run the
production-deploy.ymlworkflow manually from the GitHub Actions tab, picking aversion-bump(patch/minor/major).
Pull requests trigger the
test-lint.ymlworkflow, which runs the test suite only — no deploy.Note
The build and deploy steps are implemented as reusable workflows maintained in lambda-feedback/evaluation-function-workflows.
-
-
Once the deploy workflow has run, the platform hosts your function at a public URL. You can find it in the Admin Panel after registering the function (next step), and test it with any request client (
curl, Insomnia, Postman).Example µEd request
curl --request POST \ --url https://<your-function-url>/evaluate \ --header 'Content-Type: application/json' \ --header 'X-Api-Version: 0.1.0' \ --data '{ "submission": { "type": "MATH", "content": { "expression": "x + x" } }, "task": { "referenceSolution": { "expression": "2*x" } } }'See the µEd API section of the specification for full request/response details. Functions still running the Legacy API instead use the
commandheader — see Legacy API. -
To make your new function available on the LambdaFeedback platform, register it via the Admin Panel by supplying its name, URL and supported response types.
Note
New evaluation functions should be registered as µEd (a standard, path-based API — see Chat Functions for a general introduction to µEd on Lambda Feedback, and mued.org for the specification). The Legacy command-header API — described in the specification — is frozen and no longer developed, but Shimmy still serves it.
More Info¶
-
General Function Specification and Behaviour
- Function philosophy including deployment strategy
- Request/Response schemas and communication spec
- Base layer (Shimmy) logic, properties and behaviour
-
lf_toolkit— server wiring,Result/Params/Preview, image uploadevaluation-function-utils— the legacy package (error reporting, cross-function client)