The open-source AI hub: over a million models, hundreds of thousands of datasets, and hosted demo Spaces.
Inference
Serving model predictions in production — routing, batching and accelerated pipelines.
16 apps, 14 skills and 2 MCP servers tagged Inference.
Apps
One API and one bill for hundreds of language models, with automatic failover between providers.
Open-source desktop app to run, serve and fine-tune text, image, video and audio models entirely on your own machine.
Open-source inference engine that runs LLMs, vision and speech models fully on-device — no API keys, no cloud calls, from Swift to Godot.
Menu-bar LLM inference server for Apple Silicon, with continuous batching and tiered KV caching that keeps local models fast enough for real coding work.
Privacy-first AI gateway putting 600+ models from 90+ providers behind one OpenAI-compatible API, with no prompt or output logging.
8-29 MB automation models that run on phones, wearables and microcontrollers — tool calling, structured extraction and embeddings, fully offline.
Low-latency speech APIs from the team behind the open-source Dia model — text-to-speech and transcription priced well below the incumbents.
Open-source local inference server that profiles your hardware, picks models that fit, and points your coding agent at them — free, private and offline.
A hosted LLM gateway from ngrok: one base URL and one key for public providers, your own keys and self-hosted models, with failover, scoped access and cost analytics.
Production inference for open and custom models — fast, autoscaling, and deployable into your own cloud.
Serverless GPUs from a Python decorator — deploy models and batch jobs with no containers or cluster to manage.
Inference, fine-tuning, and GPU clusters for open models — the full stack for teams building on open weights.
A generative media platform tuned for speed — image, video, and audio models served with very low latency.
Run thousands of open models with one API call, and deploy your own without touching Kubernetes.
Inference on custom LPU hardware, built for latency — open models served at speeds general-purpose GPUs struggle to match.
Skills
AMD's official skill for CPU inference: brings up a single vLLM + zentorch endpoint on an EPYC server, sized to the socket it runs on.
AMD's official skill for standing up a vLLM endpoint on Instinct MI300X/MI325X/MI350X/MI355X GPUs — detection, recipe lookup, launch and health check in one flow.
Skill: NVIDIA TAO RT-DETR
by NVIDIA
Trains, distils, quantises, evaluates, exports and runs inference for RT-DETR real-time object detection in NVIDIA TAO, routing training through AutoML hyperparameter search by default.
Skill: NVIDIA Holoscan SDK Setup
by NVIDIA
Inspects the host's hardware, OS, CUDA driver and existing tooling, recommends one Holoscan SDK install method with a reason, then hands off to the matching install skill.
Skill: NVIDIA NV-Reason-CXR
by NVIDIA
Runs NVIDIA's NV-Reason-CXR-3B chest X-ray reasoning model through a documented wrapper, locally on GPU or via the public Hugging Face Space, and returns the model's full reasoning trace as engineering evidence.
Skill: Replicate: Run Models
by Replicate
Replicate's official skill for running models via its API — predictions, polling, Prefer: wait, webhooks with signature validation, streaming and file handling.
Skill: Jetson LLM Serve
by NVIDIA
Picks the right vLLM or SGLang runtime for your Jetson generation and JetPack version, then produces a working OpenAI-compatible serving command.
Turns a plain-language weather question into a working Earth2Studio inference script — picks the AI forecast model, a compatible data source, an IO backend, and the step count.
NVIDIA's official skill for deploying and operating Nemotron Speech (Riva) NIMs — ASR, text-to-speech and translation, cloud-hosted or self-hosted on your own GPUs.
NVIDIA's official skill for generating and validating runnable gst-launch DeepStream pipelines from a plain-language description of the video inference you want.
Skill: Dynamo Router Starter
by NVIDIA
NVIDIA's official skill for bringing up Dynamo inference routing — pick a router mode, enable KV-aware routing, and smoke-test the frontend endpoint before claiming anything works.
Skill: Local Model Selection
by Hugging Face
Choose and run the right local model with llama.cpp and GGUF — quantisation, hardware fit, and local serving.
Skill: Transformers.js
by Hugging Face
Run machine learning models directly in the browser or Node with Transformers.js — no inference server required.
Skill: Claude API Reference
by Anthropic
Authoritative reference for the Claude API and Anthropic SDKs — model IDs, pricing, parameters, streaming, tool use, MCP, agents, prompt caching, and token counting.
MCP servers
Replicate's official MCP server: search thousands of hosted models, read their schemas, and run predictions on image, video, audio and language models from inside an agent.
MCP: Hugging Face
by Hugging Face
Hugging Face's official MCP server — search the Hub's models, datasets and Spaces, read repository files, and call thousands of Gradio applications as tools from one remote endpoint.
Related tags
Tags that appear alongside this one, ranked by how often.