LLM-readable documentation index

Score Function Against (input, expected) Pairs

Hand off to an LLM

Score a function against a list of (input, expected) pairs.

Submits a batch of (input, expected) pairs, runs the named function over each input, and returns per-pair + aggregate accuracy metrics comparing the function's actual output to the provided expected JSON.

Scoring runs asynchronously. The response carries a scoreRunID; poll GET /v3/eval/score/{scoreRunID} until status is one of completed, error, or cancelled.

This request says only what to extract. How the output is compared against the expected value happens on the GET, recomputed from stored JSON each time.

POST
/v3/eval/score
x-api-key<token>

Authenticate using API Key in request header

In: header

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

curl -X POST "https://api.bem.ai/v3/eval/score" \  -H "Content-Type: application/json" \  -d '{    "functionName": "string"  }'
{
  "scoreRunID": "evalrun_2a8f...",
  "status": "pending"
}
{
  "message": "string",
  "code": 0,
  "details": {}
}

See also