EvalLeaderboardRow
Example Usage
typescript
import { EvalLeaderboardRow } from "@meetkai/mka1/models/components";
let value: EvalLeaderboardRow = {
agentName: "<value>",
agentVersion: "<value>",
agentEffort: "<value>",
model: "Colorado",
accuracy: 9924.22,
accuracyCi: null,
trials: 455029,
scoredTrials: 579622,
failedTrials: 294130,
costUsd: 1104.59,
runCount: 498328,
runIds: [
"<value 1>",
"<value 2>",
],
lastRunAt: 441655,
};Fields
| Field | Type | Required | Description |
|---|---|---|---|
agentName | string | ✔️ | N/A |
agentVersion | string | ✔️ | Version from the row's newest run. Not part of the row identity, so older runs in this row may have reported a different one. |
agentEffort | string | ✔️ | N/A |
model | string | ✔️ | N/A |
accuracy | number | ✔️ | Mean of metric over the row's scored trials. Null when no trial carried a numeric value for that metric — distinct from 0, which means every trial scored zero. |
accuracyCi | components.AccuracyCi | ✔️ | 95% Wald interval, p ± 1.96·sqrt(p(1-p)/n), clamped to [0, 1]. Derived per request rather than stored, so it cannot drift when a failed trial is rerun. It assumes metric is a 0/1 outcome (the harness reward); on a continuous metric the mean is still right but this interval is not, so read it only for pass/fail metrics. Null when accuracy is. |
trials | number | ✔️ | Every sample in the row's runs, whatever its status. |
scoredTrials | number | ✔️ | Trials carrying a numeric metric — the denominator behind accuracy and accuracy_ci. |
failedTrials | number | ✔️ | Trials that ended in the failed status, i.e. produced no score at all. An agent timeout is NOT one of these: the harness still verifies and scores that trial. |
costUsd | number | ✔️ | Summed trial spend, falling back to the runs' own totals when trials carry no cost. Null when neither reported any. |
runCount | number | ✔️ | N/A |
runIds | string[] | ✔️ | The runs behind this row, largest (most trials) first, ties newest first — element 0 is the run to open when the reader drills into this row, so a full benchmark job outranks a newer one-task smoke run. |
lastRunAt | number | ✔️ | Creation time of the newest run in the row, unix seconds. |