Skip to content

Inference

Serving model predictions in production — routing, batching and accelerated pipelines.

11 apps, 9 skills and 1 MCP server tagged Inference.

Apps

All apps

The open-source AI hub: over a million models, hundreds of thousands of datasets, and hosted demo Spaces.

Coding & DevelopmentFreemium
Featured

One API and one bill for hundreds of language models, with automatic failover between providers.

Coding & DevelopmentFreemium

App: Unsloth

Unsloth AI

Featured

Open-source desktop app to run, serve and fine-tune text, image, video and audio models entirely on your own machine.

Coding & DevelopmentFree

Open-source inference engine that runs LLMs, vision and speech models fully on-device — no API keys, no cloud calls, from Swift to Godot.

Coding & DevelopmentFree

A hosted LLM gateway from ngrok: one base URL and one key for public providers, your own keys and self-hosted models, with failover, scoped access and cost analytics.

Coding & DevelopmentPaid

Production inference for open and custom models — fast, autoscaling, and deployable into your own cloud.

Coding & DevelopmentFreemium

Serverless GPUs from a Python decorator — deploy models and batch jobs with no containers or cluster to manage.

Coding & DevelopmentFreemium

Inference, fine-tuning, and GPU clusters for open models — the full stack for teams building on open weights.

Coding & DevelopmentFreemium

A generative media platform tuned for speed — image, video, and audio models served with very low latency.

Image & DesignFreemium

Run thousands of open models with one API call, and deploy your own without touching Kubernetes.

Coding & DevelopmentFreemium

Inference on custom LPU hardware, built for latency — open models served at speeds general-purpose GPUs struggle to match.

Coding & DevelopmentFreemium

Skills

All skills
New

Replicate's official skill for running models via its API — predictions, polling, Prefer: wait, webhooks with signature validation, streaming and file handling.

New

Picks the right vLLM or SGLang runtime for your Jetson generation and JetPack version, then produces a working OpenAI-compatible serving command.

2 views

NVIDIA's official skill for bringing up Dynamo inference routing — pick a router mode, enable KV-aware routing, and smoke-test the frontend endpoint before claiming anything works.

5 views

Skill: Transformers.js

by Hugging Face

Run machine learning models directly in the browser or Node with Transformers.js — no inference server required.

6 views

Authoritative reference for the Claude API and Anthropic SDKs — model IDs, pricing, parameters, streaming, tool use, MCP, agents, prompt caching, and token counting.

MCP servers

All MCP servers
Featured

Replicate's official MCP server: search thousands of hosted models, read their schemas, and run predictions on image, video, audio and language models from inside an agent.

Related tags

Tags that appear alongside this one, ranked by how often.

All tags