EvalSampleObject
Example Usage
typescript
import { EvalSampleObject } from "@meetkai/mka1/models/components";
let value: EvalSampleObject = {
id: "<id>",
object: "eval.sample",
runId: "<id>",
taskId: "<id>",
model: "Taurus",
sampleIndex: 823042,
status: "running",
datasetRow: {},
prompt: "<value>",
target: "<value>",
responseId: null,
outputText: "<value>",
reasoning: null,
extractedOutput: null,
scores: {},
judge: {
"key": "<value>",
"key1": "<value>",
},
error: {
"key": "<value>",
},
costUsd: 6761.6,
inputTokens: 809932,
outputTokens: null,
cachedTokens: 573425,
phaseTimings: {},
createdAt: 663201,
startedAt: 954309,
completedAt: 947537,
};Fields
| Field | Type | Required | Description |
|---|---|---|---|
id | string | ✔️ | N/A |
object | "eval.sample" | ✔️ | N/A |
runId | string | ✔️ | N/A |
taskId | string | ✔️ | N/A |
model | string | ✔️ | N/A |
sampleIndex | number | ✔️ | N/A |
status | components.EvalSampleStatus | ✔️ | N/A |
datasetRow | Record<string, any> | ✔️ | N/A |
prompt | string | ✔️ | N/A |
target | string | ✔️ | N/A |
responseId | string | ✔️ | N/A |
outputText | string | ✔️ | N/A |
reasoning | string | ✔️ | Model reasoning ('thinking') captured from the response's reasoning items, kept separate from output_text so graders and extraction never see it. Truncated at 64 KB of UTF-8 with a marker appended. Null when the model returned none, when the provider only sent encrypted reasoning, and for transcription tasks. |
extractedOutput | string | ✔️ | N/A |
scores | Record<string, any> | ✔️ | N/A |
judge | Record<string, any> | ✔️ | N/A |
error | Record<string, any> | ✔️ | The exception that ended this sample, if any. For a harness trial this also carries an agent timeout — which is still a scored trial, so do not read a non-null error here as an unscored one; check status and scores. |
costUsd | number | ✔️ | Spend for this one trial in USD. Null outside harness runs, and when the harness reported no cost. |
inputTokens | number | ✔️ | Prompt tokens for the whole trial session. Includes cached_tokens. |
outputTokens | number | ✔️ | Completion tokens for the whole trial session. |
cachedTokens | number | ✔️ | The share of input_tokens served from the provider's prompt cache. Reported separately because it is what makes two agents at the same token count cost an order of magnitude apart. |
phaseTimings | Record<string, any> | ✔️ | Per-phase wall clock for a harness trial: environment_setup, agent_setup, agent_execution, verifier, each { started_at, finished_at } as ISO 8601 timestamps exactly as the harness recorded them. A phase the harness did not time is absent rather than null. Drives the trial timing bar. Null outside harness runs. |
createdAt | number | ✔️ | N/A |
startedAt | number | ✔️ | N/A |
completedAt | number | ✔️ | N/A |