Skip to content
Sentence Transformers Trainer

Sentence Transformers TrainerSkill

Added to Onei
1 views
Apache-2.0
Repository Docs

Summary

Train bi-encoders, rerankers, SPLADE sparse encoders and ColBERT multi-vector models with the loss, evaluator and template each type actually needs.

Features

  • Routes to the right path for SentenceTransformer, CrossEncoder, SparseEncoder or MultiVectorEncoder
  • Per-type production training templates rather than synthesised scripts
  • Loss-to-data-shape mapping and evaluator-to-task mapping per model type
  • Documents silent failure modes: NO_DUPLICATES for MNRL, Identity() for non-BCE CrossEncoder losses
  • Hard-negative mining, distillation, LoRA and Matryoshka embeddings
  • Precision rules (fp32 load + bf16/fp16 autocast) and TrainingArguments guidance
  • VRAM sizing with multi-GPU, FSDP and DeepSpeed, plus Hugging Face Jobs execution
  • Symptom-indexed troubleshooting for metric stagnation and Hub push failures

Install This Skill

Add this skill to your favorite AI agent in a few steps.

Any AI agent

This skill is plain instructions — it works with any assistant that accepts custom instructions or system prompts.

  1. Copy the skill content with the button below.
  2. Paste it into your agent's instruction file or system prompt (for example AGENTS.md, .cursorrules, or a custom instructions field).
  3. Ask the agent to apply the skill whenever the task matches.

Skill Content

Markdown Content

Copy this content and use it with your preferred AI agent

---
name: train-sentence-transformers
description: Train or fine-tune sentence-transformers models across `SentenceTransformer` (bi-encoder, dense or static embedding model for retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal), `CrossEncoder` (reranker, pair scoring for two-stage retrieval / pair classification), `SparseEncoder` (SPLADE, sparse embedding model for learned-sparse retrieval), and `MultiVectorEncoder` (ColBERT / late-interaction, per-token embeddings scored with MaxSim). Covers loss selection, hard-negative mining, evaluators, distillation, LoRA, Matryoshka, and Hugging Face Hub publishing. Use for any sentence-transformers training task.
---

# Train a sentence-transformers Model

**This SKILL.md is a router, not a manual.** It tells you which references and example scripts to load for your task. The actual content (recommended losses, evaluators, training-script structure, model selection, training-arg knobs, troubleshooting) lives in `references/` and `scripts/`.

**Do not synthesize a training script from this file alone.** Open the per-type production template (`scripts/train_<type>_example.py`) and copy it as your starting point. The templates contain load-bearing scaffolding (autocast helper, model-card class, logger silencing list, `force=True`, `seed`, TF32, version-compatible imports, named-evaluator metric handling) that prior agent runs have repeatedly missed when rolling their own from a synthesized snippet.

## 1. Identify the model type

| Tag | Class | What it does | When to pick |
|---|---|---|---|
| **[SentenceTransformer]** | `SentenceTransformer` (bi-encoder) | Maps each input to a fixed-dim dense vector | Retrieval, similarity, clustering, classification, paraphrase mining, dedup |
| **[CrossEncoder]** | `CrossEncoder` (reranker) | Scores `(query, passage)` pairs jointly | Two-stage retrieval (rerank top-100 from bi-encoder), pair classification |
| **[SparseEncoder]** | `SparseEncoder` (SPLADE) | Sparse vectors over the vocabulary | Learned-sparse retrieval, inverted-index backends (Elasticsearch / OpenSearch / Lucene) |
| **[MultiVectorEncoder]** | `MultiVectorEncoder` (ColBERT) | One embedding per token, scored with MaxSim | Late-interaction retrieval, recall gains over bi-encoders at higher storage cost, multimodal (ColPali / ColQwen2) |

Tiebreakers when the request is ambiguous: "embedding model" / "vector search" / "similarity" → **[SentenceTransformer]**. "rerank" / "ranker" / "two-stage" → **[CrossEncoder]**. "SPLADE" / "sparse" / "inverted index" → **[SparseEncoder]**. "ColBERT" / "late interaction" / "multi-vector" / "MaxSim" / "ColPali" / "ColQwen" → **[MultiVectorEncoder]**. If still unclear, ask.

## 2. Required reading

**Read these in full before writing any code. Do not triage by perceived relevance.**

### Per-type: always required

**[SentenceTransformer]**
- `references/losses_sentence_transformer.md`: loss-to-data-shape mapping, `BatchSamplers.NO_DUPLICATES` requirement for MNRL-family, `Cached*` ↔ `gradient_checkpointing` incompatibility.
- `references/evaluators_sentence_transformer.md`: evaluator-to-task mapping, `metric_for_best_model` key construction (named vs unnamed), per-evaluator `primary_metric` values.
- `references/model_architectures.md`: encoder vs decoder vs static vs Router pipelines, pooling rules (mean / cls / lasttoken), auto-mean-pooling behavior for fresh-start MLM bases.
- `scripts/train_sentence_transformer_example.py`: production template. Copy this as your starting point.

**[CrossEncoder]**
- `references/losses_cross_encoder.md`: pointwise / pairwise / listwise / distillation, `pos_weight` derivation, `activation_fn=Identity()` mandatory for non-BCE losses (silent eval-rank collapse otherwise).
- `references/evaluators_cross_encoder.md`: `CrossEncoderRerankingEvaluator` recipe, named-evaluator key format `eval_{name}_{primary_metric}`.
- `scripts/train_cross_encoder_example.py`: production template. Copy this as your starting point.

**[SparseEncoder]**
- `references/losses_sparse_encoder.md`: `SpladeLoss` wrapper requirement, FLOPS regularizer weights, smoke-test active-dim ramp behavior.
- `references/evaluators_sparse_encoder.md`: `SparseNanoBEIREvaluator` (English-only) and the in-domain alternative, `eval_{name}_{primary_metric}` key format.
- `scripts/train_sparse_encoder_example.py`: production template. Copy this as your starting point.

**[MultiVectorEncoder]**
- `references/losses_multi_vector_encoder.md`: MaxSim scoring, scale choice per scoring mode (`scale=1.0` for MaxSim, roughly the average query length for MeanMaxSim), MNRL / CachedMNRL / MarginMSE / DistillKLDiv, XTR-vs-ColBERT scoring, CachedMNRL ↔ `gradient_checkpointing` incompatibility.
- `references/evaluators_multi_vector_encoder.md`: `MultiVectorNanoBEIREvaluator` (English-only) and the in-domain alternative, `eval_NanoBEIR_mean_maxsim_ndcg@10` key format, distillation-eval spearman variant.
- `scripts/train_multi_vector_encoder_example.py`: production template. Copy this as your starting point.

### Cross-cutting: always required (regardless of task)

- `references/training_args.md`: `TrainingArguments` knobs, precision rules (load fp32 + autocast bf16/fp16, never `torch_dtype=bfloat16`), `warmup_steps` (float) vs deprecated `warmup_ratio`, `save_steps` must be a multiple of `eval_steps` for `load_best_model_at_end`, schedulers, HPO, tracker, resume, hub-push variants.
- `references/dataset_formats.md`: column-matching rules (label name auto-detection, column-order-not-name), reshaping recipes, hard-negative mining options.
- `references/base_model_selection.md`: discovery commands, per-type model namespaces, ModernBERT-family `max_seq_length=8192` trap, `datasets >= 4` script-loader rejection, non-English starting-point shortcuts.
- `references/troubleshooting.md`: symptom-indexed failure recipes. Skim the section headings on every run, even a healthy one. The "Metrics don't improve" and "Hub push fails" entries cover bugs that bite frequently and are cheaper to recognize before they fire than to debug after.

### Cross-cutting: load when applicable

- `references/hardware_guide.md`: VRAM sizing, multi-GPU, FSDP / DeepSpeed, HF Jobs flavors. Required for >24GB models, multi-GPU, or HF Jobs runs.
- `references/hf_jobs_execution.md`: required when running on HF Jobs.
- `references/prompts_and_instructions.md`: required when using prompt-tuned bases (E5, BGE, GTE, Qwen3-Embedding, Instructor, Nomic, etc.) or adding `query: ` / `passage: ` style prefixes.

### Variant scripts (open when the task matches)
- **[SentenceTransformer]** `scripts/train_sentence_transformer_<matryoshka|multi_dataset|with_lora|distillation|make_multilingual|static_embedding>_example.py`.
- **[CrossEncoder]** `scripts/train_cross_encoder_<distillation|listwise>_example.py`.
- **[SparseEncoder]** `scripts/train_sparse_encoder_distillation_example.py`.
- Hard-negative mining CLI: `scripts/mine_hard_negatives.py`.

## 3. Defaults

Override only if the user specifies otherwise:
- **Local execution.** Pitch HF Jobs only if local hardware can't fit the job.
- **Single run.** After it completes, propose experimentation if the user would benefit (weak/marginal verdict, "see how high you can push it" framing, etc.). Iteration rules in `references/training_args.md` (Experimentation section).
- **Public Hub push at end-of-run, wrapped in try-except.** On HF Jobs (ephemeral env) ALSO enable in-trainer push (`push_to_hub=True` + `hub_strategy="every_save"`). Details in `references/hf_jobs_execution.md`.

## 4. Constraints the produced script must satisfy

These are non-negotiable contracts. Implementation lives in the production templates and references. Do not reinvent.

- Capture the pre-training evaluator score as `baseline_eval` **before** `trainer.train()`.
- Emit a single end-of-run line: `VERDICT: WIN|MARGINAL|REGRESSION | score=... | baseline=... | delta=...`. A monitor scrapes for this.
- Silence `httpx`, `httpcore`, `huggingface_hub`, `urllib3`, `filelock`, `fsspec` to WARNING (otherwise HF download URLs flood the agent's context).
- Tee logs to `logs/{RUN_NAME}.log`.
- End with `model.push_to_hub(...)` wrapped in `try/except`.
- Smoke-test before any long run (`max_steps=1` + tiny dataset slice). The production templates show one common pattern (`SMOKE_TEST` env var).
- **[CrossEncoder]** Include `EarlyStoppingCallback(patience>=3)`. CE rerankers often peak mid-training and regress.
- **[SparseEncoder]** Log `query_active_dims` / `corpus_active_dims` on the verdict line. High nDCG with collapsed sparsity is not a win. The keys come back name-prefixed (e.g. `..._query_active_dims`). Use suffix matching to pluck them. See the SPARSE production template for the exact pattern.
- **[MultiVectorEncoder]** Match `scale` to the scoring mode on any MNRL-family loss: near `1.0` for unnormalized MaxSim (do not copy `scale=20.0` from bi-encoder MNRL), roughly the average query length with length-normalized MeanMaxSim, since each score is divided by its query's token count. `XTRScores` is a train-only `similarity_fct`: the evaluators reject it, so evaluation always scores with MaxSim, including for XTR-trained models.

## 5. Workflow

1. Identify the model type (§1). Ask if ambiguous.
2. Load the §2 required-reading files for that type.
3. Open `scripts/train_<type>_example.py` and copy it as your starting point.
4. Replace `MODEL_NAME`, `DATASET_NAME`, `RUN_NAME`, the loss, and the evaluator with the user's task. Cross-check loss/data-shape match against `references/losses_<type>.md`. Cross-check the `metric_for_best_model` key against `references/evaluators_<type>.md` (named evaluators format the key as `eval_{name}_{primary_metric}`).
5. Smoke-test (`max_steps=1`).
6. Run.
7. After the run, append to `logs/experiments.md` and propose iteration if the verdict is weak/marginal.

## Prerequisites

```bash
pip install "sentence-transformers[train]>=5.0"        # add [train,image] / [audio] / [video] for [SentenceTransformer] multimodal
                                                       # [MultiVectorEncoder] requires >=6.0
pip install trackio                                    # optional tracker (or wandb / tensorboard / mlflow)
hf auth login                                          # or set HF_TOKEN with write scope (for Hub push)
```

GPU strongly recommended. CPU works only for demos and `[SentenceTransformer]` `StaticEmbedding`.

Usage Instructions

Learn how to use this skill with different AI agents.

Claude Desktop

Install the Hugging Face marketplace once, then add this skill with the hf CLI:

/plugin marketplace add huggingface/skills
/plugin install hf-cli@huggingface/skills
hf skills add train-sentence-transformers

Description

Hugging Face's sentence-transformers training skill is deliberately built as a router rather than a manual, and that design choice is the point. It first pins down which of four model types you are training — SentenceTransformer (bi-encoder, dense or static embeddings for retrieval, clustering, dedup and paraphrase mining), CrossEncoder (reranker scoring query/passage pairs for two-stage retrieval), SparseEncoder (SPLADE-style sparse vectors for inverted-index backends such as Elasticsearch or OpenSearch), or MultiVectorEncoder (ColBERT-style per-token embeddings scored with MaxSim) — and then tells the agent exactly which reference files and production template to open.

It explicitly forbids synthesising a training script from the skill file alone, and gives the reason: the per-type templates carry load-bearing scaffolding that agents repeatedly reinvent badly — the autocast helper, model-card class, logger-silencing list, force=True, seeding, TF32 handling, version-compatible imports and named-evaluator metric keys.

The reference material is where the hard-won detail lives, and it is unusually specific about silent failure modes: BatchSamplers.NO_DUPLICATES is required for the MNRL loss family; Cached* losses are incompatible with gradient checkpointing; activation_fn=Identity() is mandatory for non-BCE CrossEncoder losses or evaluation rank collapses without an error; save_steps must be a multiple of eval_steps for load_best_model_at_end; models should be loaded in fp32 with bf16/fp16 autocast rather than torch_dtype=bfloat16; and the ModernBERT family carries a max_seq_length=8192 trap.

Also covered: hard-negative mining, evaluator-to-task mapping, distillation, LoRA, Matryoshka embeddings, base-model selection per type, VRAM sizing with multi-GPU/FSDP/DeepSpeed guidance, running on Hugging Face Jobs, and publishing the finished model to the Hub.

Related Skills

CodeQL, Semgrep and SARIF static-analysis toolkit from Trail of Bits: taint tracking, fast pattern scans and merged, deduplicated security findings for coding agents.

5 views
New

Microsoft's official Playwright skill — drives a real browser from the command line using accessibility snapshots and element refs, and plans, generates and heals Playwright tests.

7 views 1 copies

Skill: Google Maps Platform

by Google Maps Platform

New

Google's official agent skill for writing production Maps Platform code — grounded in freshly fetched docs, with a demo key path that needs no billing account.

9 views
Browse all skills →