From Cerebras Inference · · 1 min
Inside the Cerebras Inference Model Catalog
Discover how Cerebras delivers ultra-fast, unpruned open-source AI models with high precision.
In brief
Discover how Cerebras delivers ultra-fast, unpruned open-source AI models with high precision. Cerebras delivers high-speed inference on unpruned, original model architectures with selective quantization for maximum precision. Originally reported by Cerebras Inference.
Flexible inference endpoint tiers
Cerebras offers public endpoints for free trials and pay-as-you-go usage, alongside dedicated endpoints built for custom throughput and production SLAs.
“Models on Cerebras public endpoints are available on the free trial and pay-as-you-go tiers, subject to rate limits and pricing.”
Ultra-fast open model performance
Publicly hosted open-source models deliver extreme speeds, reaching up to 3,000 tokens per second for 120-billion parameter architectures.
Strict zero-pruning public policy
Public API endpoints host unpruned model versions to preserve original quality. Experimental research models like REAP live strictly on Hugging Face.
“All models served through our public endpoints are the original, unpruned versions.”
Selective weight-only quantization
Weights use selective quantization in storage, while activations, attention, and KV cache stay unquantized in full precision during operation.
“The activations, attention, and kv cache remain in full precision and unquantized.”
Guaranteed model integrity
Model architectures are never modified behind the scenes. Any future compressed variants will be offered under transparent, distinct endpoint names.
“We are committed to serving the original models for all existing endpoints, without modification.”
What it means
Cerebras delivers high-speed inference on unpruned, original model architectures with selective quantization for maximum precision.





