Skip to content

Inference

Serving model predictions in production — routing, batching and accelerated pipelines.

16 apps, 14 skills and 2 MCP servers tagged Inference.

Apps

All apps

The open-source AI hub: over a million models, hundreds of thousands of datasets, and hosted demo Spaces.

Coding & DevelopmentFreemium
Featured

One API and one bill for hundreds of language models, with automatic failover between providers.

Coding & DevelopmentFreemium

App: Unsloth

Unsloth AI

Featured

Open-source desktop app to run, serve and fine-tune text, image, video and audio models entirely on your own machine.

Coding & DevelopmentFree
Featured

Open-source inference engine that runs LLMs, vision and speech models fully on-device — no API keys, no cloud calls, from Swift to Godot.

Coding & DevelopmentFree

App: oMLX

Jun Kim

Featured

Menu-bar LLM inference server for Apple Silicon, with continuous batching and tiered KV caching that keeps local models fast enough for real coding work.

Coding & DevelopmentFree

Privacy-first AI gateway putting 600+ models from 90+ providers behind one OpenAI-compatible API, with no prompt or output logging.

Coding & DevelopmentPaid

App: Cactus

Cactus Compute, Inc.

New

8-29 MB automation models that run on phones, wearables and microcontrollers — tool calling, structured extraction and embeddings, fully offline.

Coding & DevelopmentFreemium

Low-latency speech APIs from the team behind the open-source Dia model — text-to-speech and transcription priced well below the incumbents.

Audio & VoicePaid

Open-source local inference server that profiles your hardware, picks models that fit, and points your coding agent at them — free, private and offline.

Coding & DevelopmentFree

A hosted LLM gateway from ngrok: one base URL and one key for public providers, your own keys and self-hosted models, with failover, scoped access and cost analytics.

Coding & DevelopmentPaid

Production inference for open and custom models — fast, autoscaling, and deployable into your own cloud.

Coding & DevelopmentFreemium

Serverless GPUs from a Python decorator — deploy models and batch jobs with no containers or cluster to manage.

Coding & DevelopmentFreemium

Inference, fine-tuning, and GPU clusters for open models — the full stack for teams building on open weights.

Coding & DevelopmentFreemium

A generative media platform tuned for speed — image, video, and audio models served with very low latency.

Image & DesignFreemium

Run thousands of open models with one API call, and deploy your own without touching Kubernetes.

Coding & DevelopmentFreemium

Inference on custom LPU hardware, built for latency — open models served at speeds general-purpose GPUs struggle to match.

Coding & DevelopmentFreemium

Skills

All skills

AMD's official skill for standing up a vLLM endpoint on Instinct MI300X/MI325X/MI350X/MI355X GPUs — detection, recipe lookup, launch and health check in one flow.

2 views

Trains, distils, quantises, evaluates, exports and runs inference for RT-DETR real-time object detection in NVIDIA TAO, routing training through AutoML hyperparameter search by default.

7 views

Inspects the host's hardware, OS, CUDA driver and existing tooling, recommends one Holoscan SDK install method with a reason, then hands off to the matching install skill.

5 views

Runs NVIDIA's NV-Reason-CXR-3B chest X-ray reasoning model through a documented wrapper, locally on GPU or via the public Hugging Face Space, and returns the model's full reasoning trace as engineering evidence.

6 views

Replicate's official skill for running models via its API — predictions, polling, Prefer: wait, webhooks with signature validation, streaming and file handling.

8 views

Picks the right vLLM or SGLang runtime for your Jetson generation and JetPack version, then produces a working OpenAI-compatible serving command.

12 views

NVIDIA's official skill for bringing up Dynamo inference routing — pick a router mode, enable KV-aware routing, and smoke-test the frontend endpoint before claiming anything works.

12 views

Skill: Transformers.js

by Hugging Face

Run machine learning models directly in the browser or Node with Transformers.js — no inference server required.

12 views

Authoritative reference for the Claude API and Anthropic SDKs — model IDs, pricing, parameters, streaming, tool use, MCP, agents, prompt caching, and token counting.

MCP servers

All MCP servers
Featured

Replicate's official MCP server: search thousands of hosted models, read their schemas, and run predictions on image, video, audio and language models from inside an agent.

MCP: Hugging Face

by Hugging Face

Hugging Face's official MCP server — search the Hub's models, datasets and Spaces, read repository files, and call thousands of Gradio applications as tools from one remote endpoint.

Related tags

Tags that appear alongside this one, ranked by how often.

All tags