Skip to content

EvalSampleObject ​

Example Usage ​

typescript
import { EvalSampleObject } from "@meetkai/mka1/models/components";

let value: EvalSampleObject = {
  id: "<id>",
  object: "eval.sample",
  runId: "<id>",
  taskId: "<id>",
  model: "Taurus",
  sampleIndex: 823042,
  status: "running",
  datasetRow: {},
  prompt: "<value>",
  target: "<value>",
  responseId: null,
  outputText: "<value>",
  reasoning: null,
  extractedOutput: null,
  scores: {},
  judge: {
    "key": "<value>",
    "key1": "<value>",
  },
  error: {
    "key": "<value>",
  },
  costUsd: 6761.6,
  inputTokens: 809932,
  outputTokens: null,
  cachedTokens: 573425,
  phaseTimings: {},
  createdAt: 663201,
  startedAt: 954309,
  completedAt: 947537,
};

Fields ​

FieldTypeRequiredDescription
idstring✔️N/A
object"eval.sample"✔️N/A
runIdstring✔️N/A
taskIdstring✔️N/A
modelstring✔️N/A
sampleIndexnumber✔️N/A
statuscomponents.EvalSampleStatus✔️N/A
datasetRowRecord<string, any>✔️N/A
promptstring✔️N/A
targetstring✔️N/A
responseIdstring✔️N/A
outputTextstring✔️N/A
reasoningstring✔️Model reasoning ('thinking') captured from the response's reasoning items, kept separate from output_text so graders and extraction never see it. Truncated at 64 KB of UTF-8 with a marker appended. Null when the model returned none, when the provider only sent encrypted reasoning, and for transcription tasks.
extractedOutputstring✔️N/A
scoresRecord<string, any>✔️N/A
judgeRecord<string, any>✔️N/A
errorRecord<string, any>✔️The exception that ended this sample, if any. For a harness trial this also carries an agent timeout — which is still a scored trial, so do not read a non-null error here as an unscored one; check status and scores.
costUsdnumber✔️Spend for this one trial in USD. Null outside harness runs, and when the harness reported no cost.
inputTokensnumber✔️Prompt tokens for the whole trial session. Includes cached_tokens.
outputTokensnumber✔️Completion tokens for the whole trial session.
cachedTokensnumber✔️The share of input_tokens served from the provider's prompt cache. Reported separately because it is what makes two agents at the same token count cost an order of magnitude apart.
phaseTimingsRecord<string, any>✔️Per-phase wall clock for a harness trial: environment_setup, agent_setup, agent_execution, verifier, each { started_at, finished_at } as ISO 8601 timestamps exactly as the harness recorded them. A phase the harness did not time is absent rather than null. Drives the trial timing bar. Null outside harness runs.
createdAtnumber✔️N/A
startedAtnumber✔️N/A
completedAtnumber✔️N/A