The open-source AI hub: over a million models, hundreds of thousands of datasets, and hosted demo Spaces.
Inference
Serving model predictions in production — routing, batching and accelerated pipelines.
10 apps and 2 skills tagged Inference.
Apps
One API and one bill for hundreds of language models, with automatic failover between providers.
Open-source desktop app to run, serve and fine-tune text, image, video and audio models entirely on your own machine.
Inference on custom LPU hardware, built for latency — open models served at speeds general-purpose GPUs struggle to match.
Run thousands of open models with one API call, and deploy your own without touching Kubernetes.
A generative media platform tuned for speed — image, video, and audio models served with very low latency.
Inference, fine-tuning, and GPU clusters for open models — the full stack for teams building on open weights.
Serverless GPUs from a Python decorator — deploy models and batch jobs with no containers or cluster to manage.
Production inference for open and custom models — fast, autoscaling, and deployable into your own cloud.
A hosted LLM gateway from ngrok: one base URL and one key for public providers, your own keys and self-hosted models, with failover, scoped access and cost analytics.
Skills
NVIDIA's official skill for generating and validating runnable gst-launch DeepStream pipelines from a plain-language description of the video inference you want.
NVIDIA's official skill for bringing up Dynamo inference routing — pick a router mode, enable KV-aware routing, and smoke-test the frontend endpoint before claiming anything works.
Related tags
Tags that appear alongside this one, ranked by how often.