Skip to content

Helper Packages

A toolkit implements Shimmy's worker interface so your function only has to provide comparison logic. Whether one is available depends on the language and the base layer:

Language Toolkit Used by Provides
Python lf_toolkit — repo toolkit-python Functions on the Shimmy python base image Server wiring, Result / Params / Preview, image upload
Python (legacy) evaluation-function-utils Functions on the older AWS Lambda base layer EvaluationException, cross-function client
Wolfram Language toolkit-wolfram Functions on the Shimmy wolfram base image ServeEvaluationFunction, transport wiring, error catching
Lean, or any other language none yet (can be provided on request) Functions on the lean / scratch base images — the function talks to Shimmy directly over the file interface

The Python and Wolfram toolkits are loaded and wired up automatically by their base image. A Lean or scratch function has no toolkit today: it reads the request file and writes the response file itself — see Other Languages. If you are building functions in a language without a toolkit and would benefit from one, the Lambda Feedback team can provide it on request — open an issue on shimmy.

lf_toolkit

lf_toolkit (repo toolkit-python) is pulled in via the boilerplate's pyproject.toml and pre-installed in the evaluation-function-base/python image.

Server wiring

evaluation_function/main.py connects your function to Shimmy:

from lf_toolkit import create_server, run
from .evaluation import evaluation_function
from .preview import preview_function


def main():
    server = create_server()
    server.eval(evaluation_function)
    server.preview(preview_function)
    run(server)


if __name__ == "__main__":
    main()

create_server() reads the EVAL_IO / EVAL_RPC_TRANSPORT environment variables that Shimmy injects and returns the right server (stdio, IPC or file). healthcheck is provided by the toolkit — it discovers and runs the *_test.py files — so you do not register it yourself.

Result, Params, Preview

from lf_toolkit.evaluation import Result, Params
from lf_toolkit.preview import Preview


def evaluation_function(response, answer, params: Params) -> Result:
    return Result(is_correct=response == answer)
  • Result — is_correct, tagged feedback (add_feedback(tag, text)), response_latex, response_simplified. Shimmy serialises it (is_correct, feedback, …).
  • Params — dict-like wrapper over the request params.
  • Preview — the value returned from preview_function (latex, sympy, feedback).

Image upload

lf_toolkit.evaluation.image_upload provides upload_image(...) and ImageUploadError for functions that return generated images.

Errors

lf_toolkit has no structured-exception class. Raising any exception from your function makes Shimmy stop the evaluation and return:

{ "error": { "message": "<repr of the exception>" } }

toolkit-wolfram

toolkit-wolfram — the "Evaluation Function Toolkit for Wolfram" — is the Wolfram-language equivalent of lf_toolkit. It is cloned into the evaluation-function-base/wolfram image at a tagged version and loaded by that image's Bootstrap.wl.

A Wolfram function repo does not call the toolkit directly. It only provides evaluate.m and preview.m defining evaluate`EvaluationFunction and preview`PreviewFunction; the base image's FUNCTION_COMMAND / FUNCTION_ARGS already point Shimmy at Bootstrap.wl, which loads the toolkit and wires them up.

For custom wiring or local testing, call ServeEvaluationFunction[EvaluationFunction, PreviewFunction] directly — it reads Shimmy's EVAL_IO / EVAL_RPC_TRANSPORT contract and dispatches to whichever transport Shimmy selected (the file interface, or an RPC transport). A Wolfram error raised by your function is caught and returned as an error response instead of crashing the worker. See the toolkit-wolfram README for the exact contract and the current list of supported transports.

evaluation-function-utils (legacy)

Note

This package is only present on the older AWS Lambda base layer. New functions on Shimmy use lf_toolkit instead.

Errors

Submodule containing custom error and exception classes, which can be properly caught by the base evaluation layer, and return more detailed and appropriate errors.

class EvaluationException

This class extends the usual python Exception, with additional functionality. It can be used to package additional fields and values to errors thrown and returned by evaluation functions.

Example

If at some point in the execution of the evaluation_function, an error is thrown:

from evaluation_function_utils.errors import EvaluationException

if isinstance(input, str):
    raise EvaluationException(
        "The input must not be a string", 
        valid_types=["int", "float", "array"],
    )

Then the output generated by the lambda function will look like:

{
  "command": "eval",
  "error": {
    "message": "The input must not be a string",
    "valid_types": [
      "int", "float", "array"
    ]
  }
}

This class contains an error_dict property, which packages the additional arguments given to the Exception instance into a JSON-serializable object. It does so in an error-safe way, also reporting serialization errors if they occur.

Client

This submodule contains a custom EvaluationFunctionClient, which can be used to call other deployed evaluation functions.

class EvaluationFunctionClient

Client wrapped around the botocore.client.Lambda, for invoking deployed evaluation functions. On initialisation, it fetches credentials from environment variables "INVOKER_KEY", "INVOKER_ID" and "INVOKER_REGION", or from an optional environment file prescrived by env_path.

Example

from evaluation_function_utils.client import EvaluationFunctionClient
client = EvaluationFunctionClient()

def evaluation_function(response, answer, params): 
    return client.invoke('isExactEqual', response, answer, params)

In this example, the evaluation_function completely offloads grading to the deployed 'isExactEqual' function.

Note: The EvaluationFunctionClient.invoke method was designed to behave exactly as if the evaluation_function function defined in the targeted deployed function was called directly. This means that if errors are encountered an EvaluationException is raised.