# Inside the Cerebras Inference Model Catalog

> Discover how Cerebras delivers ultra-fast, unpruned open-source AI models with high precision.
- Title: Inside the Cerebras Inference Model Catalog — Shortify
- Summary: Discover how Cerebras delivers ultra-fast, unpruned open-source AI models with high precision. Cerebras delivers high-speed inference on unpruned, original…
- Keywords: cerebras, ai inference, llm, quantization, open source, technology, startups, Inside, Inference, Model, Catalog, Cerebras Inference
- Source: Cerebras Inference — https://inference-docs.cerebras.ai/models/overview
- Read time: 1 min
- Topics: cerebras, ai inference, llm, quantization, open source, technology, startups
## Flexible inference endpoint tiers
Cerebras offers public endpoints for free trials and pay-as-you-go usage, alongside dedicated endpoints built for custom throughput and production SLAs.

> Models on Cerebras public endpoints are available on the free trial and pay-as-you-go tiers, subject to rate limits and pricing.
## Ultra-fast open model performance
Publicly hosted open-source models deliver extreme speeds, reaching up to 3,000 tokens per second for 120-billion parameter architectures.
## Strict zero-pruning public policy
Public API endpoints host unpruned model versions to preserve original quality. Experimental research models like REAP live strictly on Hugging Face.

> All models served through our public endpoints are the original, unpruned versions.
## Selective weight-only quantization
Weights use selective quantization in storage, while activations, attention, and KV cache stay unquantized in full precision during operation.

> The activations, attention, and kv cache remain in full precision and unquantized.
## Guaranteed model integrity
Model architectures are never modified behind the scenes. Any future compressed variants will be offered under transparent, distinct endpoint names.

> We are committed to serving the original models for all existing endpoints, without modification.
## Key takeaway

Cerebras delivers high-speed inference on unpruned, original model architectures with selective quantization for maximum precision.