Skip to content

Inference

Serving model predictions in production — routing, batching and accelerated pipelines.

10 apps and 2 skills tagged Inference.

Apps

All apps

The open-source AI hub: over a million models, hundreds of thousands of datasets, and hosted demo Spaces.

Coding & DevelopmentFreemium
Featured

One API and one bill for hundreds of language models, with automatic failover between providers.

Coding & DevelopmentFreemium

App: Unsloth

Unsloth AI

New

Open-source desktop app to run, serve and fine-tune text, image, video and audio models entirely on your own machine.

Coding & DevelopmentFree

Inference on custom LPU hardware, built for latency — open models served at speeds general-purpose GPUs struggle to match.

Coding & DevelopmentFreemium

Run thousands of open models with one API call, and deploy your own without touching Kubernetes.

Coding & DevelopmentFreemium

A generative media platform tuned for speed — image, video, and audio models served with very low latency.

Image & DesignFreemium

Inference, fine-tuning, and GPU clusters for open models — the full stack for teams building on open weights.

Coding & DevelopmentFreemium

Serverless GPUs from a Python decorator — deploy models and batch jobs with no containers or cluster to manage.

Coding & DevelopmentFreemium

Production inference for open and custom models — fast, autoscaling, and deployable into your own cloud.

Coding & DevelopmentFreemium

A hosted LLM gateway from ngrok: one base URL and one key for public providers, your own keys and self-hosted models, with failover, scoped access and cost analytics.

Coding & DevelopmentPaid

Skills

All skills

NVIDIA's official skill for bringing up Dynamo inference routing — pick a router mode, enable KV-aware routing, and smoke-test the frontend endpoint before claiming anything works.

Related tags

Tags that appear alongside this one, ranked by how often.

All tags