The open-source AI hub: over a million models, hundreds of thousands of datasets, and hosted demo Spaces.
Inference
Serving model predictions in production — routing, batching and accelerated pipelines.
11 apps, 9 skills and 1 MCP server tagged Inference.
Apps
One API and one bill for hundreds of language models, with automatic failover between providers.
Open-source desktop app to run, serve and fine-tune text, image, video and audio models entirely on your own machine.
Open-source inference engine that runs LLMs, vision and speech models fully on-device — no API keys, no cloud calls, from Swift to Godot.
A hosted LLM gateway from ngrok: one base URL and one key for public providers, your own keys and self-hosted models, with failover, scoped access and cost analytics.
Production inference for open and custom models — fast, autoscaling, and deployable into your own cloud.
Serverless GPUs from a Python decorator — deploy models and batch jobs with no containers or cluster to manage.
Inference, fine-tuning, and GPU clusters for open models — the full stack for teams building on open weights.
A generative media platform tuned for speed — image, video, and audio models served with very low latency.
Run thousands of open models with one API call, and deploy your own without touching Kubernetes.
Inference on custom LPU hardware, built for latency — open models served at speeds general-purpose GPUs struggle to match.
Skills
Replicate's official skill for running models via its API — predictions, polling, Prefer: wait, webhooks with signature validation, streaming and file handling.
Picks the right vLLM or SGLang runtime for your Jetson generation and JetPack version, then produces a working OpenAI-compatible serving command.
Turns a plain-language weather question into a working Earth2Studio inference script — picks the AI forecast model, a compatible data source, an IO backend, and the step count.
NVIDIA's official skill for deploying and operating Nemotron Speech (Riva) NIMs — ASR, text-to-speech and translation, cloud-hosted or self-hosted on your own GPUs.
NVIDIA's official skill for generating and validating runnable gst-launch DeepStream pipelines from a plain-language description of the video inference you want.
Skill: Dynamo Router Starter
by NVIDIA
NVIDIA's official skill for bringing up Dynamo inference routing — pick a router mode, enable KV-aware routing, and smoke-test the frontend endpoint before claiming anything works.
Skill: Local Model Selection
by Hugging Face
Choose and run the right local model with llama.cpp and GGUF — quantisation, hardware fit, and local serving.
Skill: Transformers.js
by Hugging Face
Run machine learning models directly in the browser or Node with Transformers.js — no inference server required.
Skill: Claude API Reference
by Anthropic
Authoritative reference for the Claude API and Anthropic SDKs — model IDs, pricing, parameters, streaming, tool use, MCP, agents, prompt caching, and token counting.
MCP servers
Replicate's official MCP server: search thousands of hosted models, read their schemas, and run predictions on image, video, audio and language models from inside an agent.
Related tags
Tags that appear alongside this one, ranked by how often.