Lorem ipsum > Lorem ipsum
The performance, infrastructure clarity, and explicit control you need to scale
Why dedicated inference
With Dedicated Inference from CoreWeave, you bring the model and make the architectural choices that matter: GPU class, runtime, scaling, routing. CoreWeave runs the cluster, manages availability, and keeps performance and cost legible as you scale.
Lorem ipsum dolar
Give your teams the compute and flexibility to move faster—from first training run to production inference.
Who this is for
The teams that get the most value from Dedicated Inference sit between “just use an API” and “run our own Kubernetes cluster.”
AI
Teams deploying domain-specific or fine-tuned models who need custom weights support, GPU selection, and a path from training.
Enterprise
Organizations that need single-tenant GPU nodes, Availability Zone placement for data sovereignty, and predictable SLA.
Platform
ML platform teams who need to grow into distributed inference while preserving GPU choice, runtime flexibility.
CoreWeave inference paths
Three inference paths built on the award-winning CoreWeave Cloud. Move between them as your workloads evolve—without replatforming—so you always get predictable performance and infrastructure-aligned economics.

Serve custom models without clusters
Bring your own weights, choose your GPU, and deploy to a live endpoint. CoreWeave manages the infrastructure and serving layer—no cluster or Kubernetes management required.
Keep inference costs predictable
Pay per GPU hour with explicitly selected GPU classes. Use existing reserved nodes for Dedicated Inference at contracted rates without incremental fees.
Go from fine-tuning to serving
Deploy weights directly from CoreWeave AI Object Storage without moving data. Move from training to production on the same infrastructure and runtimes.
How it works
Provision a gateway, configure your deployment, send requests, observe. You make the architectural choices; CoreWeave runs the cluster.
CoreWeave runs the cluster for Dedicated Inference; you run it for CKS. With Dedicated, you configure GPU, runtime, scaling, and routing, and CoreWeave handles cluster operations, autoscaling, and availability. With CKS, you own the full Kubernetes stack: runtimes, scheduling, multi-node topology, and operational responsibility. Choose Dedicated when you want managed execution without giving up GPU and runtime choice. Choose CKS when you need to own and tune the entire stack.
Yes. Reserved nodes you already hold serve Dedicated Inference deployments at your contracted rate, with no incremental platform fee. Capacity moves between training and serving as your mix changes, so a reservation bought for a training run does not sit idle once the run ends.
Yes. A new revision rolls out behind the same endpoint: replicas on the new weights come up and pass health checks before any on the old ones are drained, and the gateway keeps routing traffic throughout. Rolling back to the previous revision works the same way.
The gateway is the tenant-isolated front door to your deployments — authentication, load balancing, and request routing across replicas. CoreWeave operates it; you configure what it optimizes for: latency, data locality, or compliance.
In CoreWeave AI Object Storage, in the same region as the deployment. Fine-tuned checkpoints, custom architectures, and open-source weights all load from there, so a model trained on CoreWeave reaches serving without the data moving.
vLLM and SGLang, with OpenAI-compatible endpoints out of the box. Both are exposed as first-class choices rather than an implementation detail, so swapping runtime does not mean rebuilding the serving stack around it.
Per-deployment metrics — latency percentiles, throughput, queue depth, GPU utilization, and error rates — are in the console and over the metrics API. They carry the same labels as the rest of your CoreWeave infrastructure, so serving lands on the dashboards you already run.
Deployments run on single-tenant GPU nodes behind a tenant-isolated gateway, with traffic encrypted in transit and weights encrypted at rest. Availability Zone placement pins where requests are served for data sovereignty, and prompts and completions are neither retained nor used for training.
Inference built
on the Essential Cloud for AI
Dedicated Inference gives you predictable performance, infrastructure-aligned economics, and explicit control.
