AMD's official meta-skill: permanently reroutes an agent's image generation, text-to-speech and speech-to-text to a local Lemonade Server instead of a paid cloud API.
Sentence Transformers TrainerSkill
Summary
Train bi-encoders, rerankers, SPLADE sparse encoders and ColBERT multi-vector models with the loss, evaluator and template each type actually needs.
Features
- Routes to the right path for SentenceTransformer, CrossEncoder, SparseEncoder or MultiVectorEncoder
- Per-type production training templates rather than synthesised scripts
- Loss-to-data-shape mapping and evaluator-to-task mapping per model type
- Documents silent failure modes: NO_DUPLICATES for MNRL, Identity() for non-BCE CrossEncoder losses
- Hard-negative mining, distillation, LoRA and Matryoshka embeddings
- Precision rules (fp32 load + bf16/fp16 autocast) and TrainingArguments guidance
- VRAM sizing with multi-GPU, FSDP and DeepSpeed, plus Hugging Face Jobs execution
- Symptom-indexed troubleshooting for metric stagnation and Hub push failures
Install This Skill
Add this skill to your favorite AI agent in a few steps.
Skill Content
Usage Instructions
Learn how to use this skill with different AI agents.
Description
Hugging Face's sentence-transformers training skill is deliberately built as a router rather than a manual, and that design choice is the point. It first pins down which of four model types you are training — SentenceTransformer (bi-encoder, dense or static embeddings for retrieval, clustering, dedup and paraphrase mining), CrossEncoder (reranker scoring query/passage pairs for two-stage retrieval), SparseEncoder (SPLADE-style sparse vectors for inverted-index backends such as Elasticsearch or OpenSearch), or MultiVectorEncoder (ColBERT-style per-token embeddings scored with MaxSim) — and then tells the agent exactly which reference files and production template to open.
It explicitly forbids synthesising a training script from the skill file alone, and gives the reason: the per-type templates carry load-bearing scaffolding that agents repeatedly reinvent badly — the autocast helper, model-card class, logger-silencing list, force=True, seeding, TF32 handling, version-compatible imports and named-evaluator metric keys.
The reference material is where the hard-won detail lives, and it is unusually specific about silent failure modes: BatchSamplers.NO_DUPLICATES is required for the MNRL loss family; Cached* losses are incompatible with gradient checkpointing; activation_fn=Identity() is mandatory for non-BCE CrossEncoder losses or evaluation rank collapses without an error; save_steps must be a multiple of eval_steps for load_best_model_at_end; models should be loaded in fp32 with bf16/fp16 autocast rather than torch_dtype=bfloat16; and the ModernBERT family carries a max_seq_length=8192 trap.
Also covered: hard-negative mining, evaluator-to-task mapping, distillation, LoRA, Matryoshka embeddings, base-model selection per type, VRAM sizing with multi-GPU/FSDP/DeepSpeed guidance, running on Hugging Face Jobs, and publishing the finished model to the Hub.
Related Skills
CodeQL, Semgrep and SARIF static-analysis toolkit from Trail of Bits: taint tracking, fast pattern scans and merged, deduplicated security findings for coding agents.
Microsoft's official Playwright skill — drives a real browser from the command line using accessibility snapshots and element refs, and plans, generates and heals Playwright tests.
Google's official agent skill for writing production Maps Platform code — grounded in freshly fetched docs, with a demo key path that needs no billing account.