Manage and run evals in the OpenAI platform.
This category contains 12 nodes.
Cancel an ongoing evaluation run.
Create the structure of an evaluation that can be used to test a model's performance. An evaluation is a set of testing criteria and the config for a data […]
Kicks off a new run for a given evaluation, specifying the data source, and what model configuration to use to test. The datasource will be validated […]
Delete an evaluation.
Delete an eval run.
Get an evaluation by ID.
Get an evaluation run by ID.
Get an evaluation run output item by ID.
Get a list of output items for an evaluation run.
Get a list of runs for an evaluation.