Helper Packages¶
A toolkit implements Shimmy's worker interface so your function only has to provide comparison logic. Whether one is available depends on the language and the base layer:
| Language | Toolkit | Used by | Provides |
|---|---|---|---|
| Python | lf_toolkit — repo toolkit-python |
Functions on the Shimmy python base image |
Server wiring, Result / Params / Preview, image upload |
| Python (legacy) | evaluation-function-utils |
Functions on the older AWS Lambda base layer | EvaluationException, cross-function client |
| Wolfram Language | toolkit-wolfram |
Functions on the Shimmy wolfram base image |
ServeEvaluationFunction, transport wiring, error catching |
| Lean, or any other language | none yet (can be provided on request) | Functions on the lean / scratch base images |
— the function talks to Shimmy directly over the file interface |
The Python and Wolfram toolkits are loaded and wired up automatically by their base image. A
Lean or scratch function has no toolkit today: it reads the request file and writes the
response file itself — see Other Languages. If you are building
functions in a language without a toolkit and would benefit from one, the Lambda Feedback team
can provide it on request — open an issue on shimmy.
lf_toolkit¶
lf_toolkit (repo toolkit-python) is
pulled in via the boilerplate's pyproject.toml and pre-installed in the
evaluation-function-base/python
image.
Server wiring¶
evaluation_function/main.py connects your function to Shimmy:
from lf_toolkit import create_server, run
from .evaluation import evaluation_function
from .preview import preview_function
def main():
server = create_server()
server.eval(evaluation_function)
server.preview(preview_function)
run(server)
if __name__ == "__main__":
main()
create_server() reads the EVAL_IO / EVAL_RPC_TRANSPORT environment variables that Shimmy
injects and returns the right server (stdio, IPC or file). healthcheck is provided by the
toolkit — it discovers and runs the *_test.py files — so you do not register it yourself.
Result, Params, Preview¶
from lf_toolkit.evaluation import Result, Params
from lf_toolkit.preview import Preview
def evaluation_function(response, answer, params: Params) -> Result:
return Result(is_correct=response == answer)
Result—is_correct, tagged feedback (add_feedback(tag, text)),response_latex,response_simplified. Shimmy serialises it (is_correct,feedback, …).Params— dict-like wrapper over the requestparams.Preview— the value returned frompreview_function(latex,sympy,feedback).
Image upload¶
lf_toolkit.evaluation.image_upload provides upload_image(...) and ImageUploadError for
functions that return generated images.
Errors¶
lf_toolkit has no structured-exception class. Raising any exception from your function
makes Shimmy stop the evaluation and return:
{ "error": { "message": "<repr of the exception>" } }
toolkit-wolfram¶
toolkit-wolfram — the "Evaluation
Function Toolkit for Wolfram" — is the Wolfram-language equivalent of lf_toolkit. It is
cloned into the
evaluation-function-base/wolfram
image at a tagged version and loaded by that image's Bootstrap.wl.
A Wolfram function repo does not call the toolkit directly. It only provides evaluate.m
and preview.m defining evaluate`EvaluationFunction and preview`PreviewFunction;
the base image's FUNCTION_COMMAND / FUNCTION_ARGS already point Shimmy at Bootstrap.wl,
which loads the toolkit and wires them up.
For custom wiring or local testing, call
ServeEvaluationFunction[EvaluationFunction, PreviewFunction] directly — it reads Shimmy's
EVAL_IO / EVAL_RPC_TRANSPORT contract and dispatches to whichever transport Shimmy
selected (the file interface, or an RPC transport). A Wolfram error raised by your function is
caught and returned as an error response instead of crashing the worker. See the
toolkit-wolfram README for the exact
contract and the current list of supported transports.
evaluation-function-utils (legacy)¶
Note
This package is only present on the older AWS Lambda base layer. New functions on Shimmy use
lf_toolkit instead.
Errors¶
Submodule containing custom error and exception classes, which can be properly caught by the base evaluation layer, and return more detailed and appropriate errors.
class EvaluationException¶
This class extends the usual python Exception, with additional functionality. It can be used to
package additional fields and values to errors thrown and returned by evaluation functions.
Example
If at some point in the execution of the evaluation_function, an error is thrown:
from evaluation_function_utils.errors import EvaluationException
if isinstance(input, str):
raise EvaluationException(
"The input must not be a string",
valid_types=["int", "float", "array"],
)
Then the output generated by the lambda function will look like:
{
"command": "eval",
"error": {
"message": "The input must not be a string",
"valid_types": [
"int", "float", "array"
]
}
}
This class contains an error_dict property, which packages the additional arguments given to the Exception instance into a JSON-serializable object. It does so in an error-safe way, also reporting serialization errors if they occur.
Client¶
This submodule contains a custom EvaluationFunctionClient, which can be used to call other deployed evaluation functions.
class EvaluationFunctionClient¶
Client wrapped around the botocore.client.Lambda, for invoking deployed evaluation functions. On initialisation, it fetches credentials from environment variables "INVOKER_KEY", "INVOKER_ID" and "INVOKER_REGION", or from an optional environment file prescrived by env_path.
Example
from evaluation_function_utils.client import EvaluationFunctionClient
client = EvaluationFunctionClient()
def evaluation_function(response, answer, params):
return client.invoke('isExactEqual', response, answer, params)
In this example, the evaluation_function completely offloads grading to the deployed 'isExactEqual' function.
Note: The EvaluationFunctionClient.invoke method was designed to behave exactly as if the evaluation_function function defined in the targeted deployed function was called directly. This means that if errors are encountered an EvaluationException is raised.