Skip to content

EvalEmbeddingConfig ​

Configuration of an embedding task: which columns hold the texts and labels, the auxiliary datasets, and the scoring protocol.

Example Usage ​

typescript
import { EvalEmbeddingConfig } from "@meetkai/mka1/models/components";

let value: EvalEmbeddingConfig = {
  kind: "classification",
};

Fields ​

FieldTypeRequiredDescription
kindcomponents.Kind✔️What the vectors are scored as. Decides the dataset layout the task expects and the metrics it reports.
textColumnstring➖Row column with the text to embed (classification, multilabel, queries and corpus documents).
text1Columnstring➖STS: first sentence of the pair.
text2Columnstring➖STS: second sentence of the pair.
scoreColumnstring➖STS: gold similarity score. Retrieval/reranking: relevance score column of the qrels rows.
labelColumnstring➖Classification: label column. Multilabel: a list of labels per row.
idColumnstring➖Retrieval/reranking: id column of query and corpus rows.
titleColumnstring➖Retrieval/reranking: optional corpus title column, prepended to the document text the way MTEB does.
queryIdColumnstring➖Retrieval/reranking: query id column of the qrels and top-ranked rows.
corpusIdColumnstring➖Retrieval/reranking: document id column of the qrels rows.
corpusIdsColumnstring➖Reranking: list of candidate document ids per query in the top-ranked rows.
trainDatasetcomponents.EvalDataset➖Dataset backing an eval task.
corpusDatasetcomponents.EvalDataset➖Dataset backing an eval task.
qrelsDatasetcomponents.EvalDataset➖Dataset backing an eval task.
topRankedDatasetcomponents.EvalDataset➖Dataset backing an eval task.
samplesPerLabelnumber➖Classification/multilabel: training rows sampled per label in each experiment (MTEB: 8).
nExperimentsnumber➖Classification/multilabel: fit-and-score repetitions averaged into the reported metrics (MTEB: 10).
seednumber➖Seed for the training-row sampling.
kValuesnumber[]➖Retrieval/reranking: cutoffs for nDCG, MAP, recall, precision and MRR.
maxCharsnumber➖Each text is cut to this many characters before it is embedded. Providers reject over-long inputs with a 4xx, which fails the sample terminally and leaves a hole in the corpus; the default 60000 sits under the gateway's own 64,000-character ceiling for its embedding models. Raise it only for a model that accepts more.
maxTextsnumber➖Fail preparation when the task would embed more rows than this (eval + train + corpus + queries). Ceiling 50000; vectors are stored per sample and scored in memory.