Open-source desktop app to run, serve and fine-tune text, image, video and audio models entirely on your own machine.
GPU
Accelerated computing on GPUs, from CUDA workloads to rented inference and training capacity.
5 apps, 10 skills and 2 MCP servers tagged GPU.
Apps
Production inference for open and custom models — fast, autoscaling, and deployable into your own cloud.
Serverless GPUs from a Python decorator — deploy models and batch jobs with no containers or cluster to manage.
Inference, fine-tuning, and GPU clusters for open models — the full stack for teams building on open weights.
Inference on custom LPU hardware, built for latency — open models served at speeds general-purpose GPUs struggle to match.
Skills
Read-only diagnostics for cluster-wide SageMaker HyperPod failures on EKS or Slurm — CloudFormation errors, EFA health checks, lifecycle scripts, capacity, dangling nodes and autoscaler conflicts.
Picks the right vLLM or SGLang runtime for your Jetson generation and JetPack version, then produces a working OpenAI-compatible serving command.
GPU-accelerated Mean-CVaR and Mean-Variance portfolio construction with NVIDIA cuOpt: scenario generation, variance-capped SOCP allocations, efficient frontiers, backtests and rebalancing.
Turns a plain-language weather question into a working Earth2Studio inference script — picks the AI forecast model, a compatible data source, an IO backend, and the step count.
NVIDIA's official skill for deploying and operating Nemotron Speech (Riva) NIMs — ASR, text-to-speech and translation, cloud-hosted or self-hosted on your own GPUs.
Skill: NVIDIA DALI Dynamic Mode
by NVIDIA
NVIDIA's official skill for DALI's imperative dynamic-mode API — write GPU data loading as ordinary Python, or migrate an existing pipeline-mode graph across.
Skill: NVIDIA CUDA-Q Onboarding
by NVIDIA
NVIDIA's official onboarding skill for CUDA-Q — installs the platform, writes your first quantum kernel, picks a GPU simulator and routes you to real QPU hardware.
Skill: NVIDIA cuDF Accelerated Computing
by NVIDIA
NVIDIA's own guidance for moving pandas workloads onto GPU DataFrames with cuDF and dask-cuDF — when it pays off, and how to keep results identical.
Skill: LLM Fine-Tuning with TRL
by Hugging Face
Fine-tune language and vision models with TRL or Unsloth — supervised, preference, and reinforcement training methods.
Skill: Hugging Face Spaces
by Hugging Face
Deploy and maintain applications on Hugging Face Spaces — SDK choice, GPU hardware, model loading, and debugging.
MCP servers
Replicate's official MCP server: search thousands of hosted models, read their schemas, and run predictions on image, video, audio and language models from inside an agent.
MCP: Runpod MCP Server
by Runpod
Runpod's official MCP server for driving GPU infrastructure — create and manage Pods, Serverless endpoints, templates and network volumes from an AI client.
Related tags
Tags that appear alongside this one, ranked by how often.