EvalEmbeddingConfig
Configuration of an embedding task: which columns hold the texts and labels, the auxiliary datasets, and the scoring protocol.
Example Usage
typescript
import { EvalEmbeddingConfig } from "@meetkai/mka1/models/components";
let value: EvalEmbeddingConfig = {
kind: "classification",
};Fields
| Field | Type | Required | Description |
|---|---|---|---|
kind | components.Kind | ✔️ | What the vectors are scored as. Decides the dataset layout the task expects and the metrics it reports. |
textColumn | string | ➖ | Row column with the text to embed (classification, multilabel, queries and corpus documents). |
text1Column | string | ➖ | STS: first sentence of the pair. |
text2Column | string | ➖ | STS: second sentence of the pair. |
scoreColumn | string | ➖ | STS: gold similarity score. Retrieval/reranking: relevance score column of the qrels rows. |
labelColumn | string | ➖ | Classification: label column. Multilabel: a list of labels per row. |
idColumn | string | ➖ | Retrieval/reranking: id column of query and corpus rows. |
titleColumn | string | ➖ | Retrieval/reranking: optional corpus title column, prepended to the document text the way MTEB does. |
queryIdColumn | string | ➖ | Retrieval/reranking: query id column of the qrels and top-ranked rows. |
corpusIdColumn | string | ➖ | Retrieval/reranking: document id column of the qrels rows. |
corpusIdsColumn | string | ➖ | Reranking: list of candidate document ids per query in the top-ranked rows. |
trainDataset | components.EvalDataset | ➖ | Dataset backing an eval task. |
corpusDataset | components.EvalDataset | ➖ | Dataset backing an eval task. |
qrelsDataset | components.EvalDataset | ➖ | Dataset backing an eval task. |
topRankedDataset | components.EvalDataset | ➖ | Dataset backing an eval task. |
samplesPerLabel | number | ➖ | Classification/multilabel: training rows sampled per label in each experiment (MTEB: 8). |
nExperiments | number | ➖ | Classification/multilabel: fit-and-score repetitions averaged into the reported metrics (MTEB: 10). |
seed | number | ➖ | Seed for the training-row sampling. |
kValues | number[] | ➖ | Retrieval/reranking: cutoffs for nDCG, MAP, recall, precision and MRR. |
maxChars | number | ➖ | Each text is cut to this many characters before it is embedded. Providers reject over-long inputs with a 4xx, which fails the sample terminally and leaves a hole in the corpus; the default 60000 sits under the gateway's own 64,000-character ceiling for its embedding models. Raise it only for a model that accepts more. |
maxTexts | number | ➖ | Fail preparation when the task would embed more rows than this (eval + train + corpus + queries). Ceiling 50000; vectors are stored per sample and scored in memory. |