- Deploy leading AI models – Access models from 9+ providers, including Meta, DeepSeek, Mistral AI, and Google, through a unified platform.
- Reduce latency and inference costs – Use built-in prompt caching to achieve 90%+ cache hit rates under sustained workloads.
- Deploy your way – Run on your infrastructure or use SambaNova’s fully managed service.
Who this guide is for
System administrators managing:
- Hardware infrastructure
- Kubernetes clusters
- Inference services (models, user groups, access control)
Required skills
- Linux system administration
- Kubernetes operations (
kubectl, Helm) - Log analysis and troubleshooting
- System credential management
See the SambaStack release notes for new features, improvements, bug fixes, and version-specific updates. For inference service features and API usage, see the Developer Guide.
Deployment options
SambaStack hosted
A fully managed cloud service from SambaNova. Build and operate high-performance inference services without managing the underlying infrastructure – SambaNova provides and maintains the AI hardware and Kubernetes environment, while you configure and manage your AI services (model selection, user groups, and users).Quickstart guide
SambaStack on-prem
Runs in your own data center, giving you full control over the AI infrastructure, Kubernetes environment, and AI services you deploy. You manage everything end-to-end – model selection, user groups, and users – to support your organization’s needs.Quickstart guide
Deployment comparison
See the on-prem architecture diagram and connection flows for a full breakdown of how requests, administration, and supporting infrastructure connect.
Explore key capabilities
Supported models and bundles
Browse available models, context lengths, and features.
Deploy a model bundle
Deploy your first model bundle on SambaStack.
Prompt caching
Cut latency and cost by caching repeated context – hit rates reach 90%+ under sustained traffic.
Authentication
Set up OIDC or LDAP authentication for your deployment.

