Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
-
Updated
Sep 9, 2026 - Go
Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
Run AI models anywhere. https://muna.ai/explore
Run ComfyUI on Modal with auto-scaling, GPU snapshots, and easy model management. Try image, video generation via ComfyUI on Modal.
Own your AI video pipeline. LTX-2.3 (22B) self-hosted on your Modal GPU via a Claude Code skill — t2v, i2v, keyframes, v2v + synced audio. ~$0.02 per 5s clip, idle = $0.
Single-user training and generation platform on Modal. Train Krea 2 LoRAs, generate stills, and turn them into video with MiniMax-H3 (sound and picture in one pass) — one deploy, one URL, no infrastructure to keep alive.
State-aware hedged requests for serverless GPU inference — return the first valid result and cancel the losers with an audited receipt.
AI-powered audio transcription, voice cloning, and image generation on Modal serverless GPUs. Real-time streaming, speaker diarization, meeting minutes, saved voice profiles, and FLUX.1 image gen — all in one service.
Queue-driven, scale-from-zero GPU inference for any Kubernetes — bursts to cross-region VMs when GPUs run dry
An open-source, BYOK YouTube thumbnail and channel analyzer powered by the TRIBE v2 neuroscience model, Next.js, and serverless GPU inference via Modal.
🔎 Search your photos by meaning — semantic photo search powered by CLIP on Runpod Flash serverless GPUs
Run Metaflow steps on serverless GPUs — no infrastructure to manage
A post-AGI framework using serverless agents to optimize Earth's carrying capacity by guiding natural evolution and implementing AI nap protocols.
High-Performance Serverless event and data processing platform
One-command PowerShell deployment of Ollama running gpt-oss:20b (or any Ollama model) on Azure Container Apps serverless GPU — nginx API-key auth proxy, scale-to-zero billing.
AI-powered genetic variant analysis platform using Stanford's Evo 2 model to predict mutation pathogenicity. Built with Next.js, FastAPI, and Serverless GPU acceleration for real-time genomic research and clinical decision support.
A cost-effective, serverless AI image generation pipeline using local n8n and Modal.com to run the uncensored FLUX.2-klein-9B model on cloud A100 GPUs.
LTX-2 video generation packaged for RunPod serverless GPU workers.
Serverless GPU cold start latency and cost benchmark for LLM inference (Modal, RunPod, Replicate, Together AI).
AI-powered medical imaging platform for DICOM ingestion, MONAI multi-organ segmentation, and MPR visualization with PHI de-identification, real-time inference streaming, and clinical PDF reporting.
Moving one app's image generation onto Runpod serverless and measuring it honestly: 388.63s to 325.24s end to end, $1.13 total spend, 0 differing pixels against the home render. Every number cites a log that a script checks.
To associate your repository with the serverless-gpu topic, visit your repo's landing page and select "manage topics."