Skip to main content
SambaStack supports a variety of models that can be deployed to both on-prem and hosted environments. Contact your system administrator to determine which models are available on your deployment. For definitions of the Preview and Production designations, see the Glossary.

Deployment options

When deploying models in SambaStack, administrators can select from various context length and batch size combinations.
  • Smaller batch sizes provide higher token throughput (tokens/second).
  • Larger batch sizes provide better concurrency for multiple users.
SambaStack supports two deployment configurations – high-interactivity and high-throughput. See High-throughput deployment for when to use each.

Supported models

You can run the following command to discover available models in your cluster:
The table below lists supported models, context lengths, batch sizes, and features. ‡ No prebuilt bundle is available for this model in SambaStack. You can use this model to create a custom bundle. In SambaStack, models are not deployed individually; they are deployed as bundles. A bundle is a packaged deployment that groups one or more models together with their associated configurations, such as batch size and sequence length. For example, deploying the Meta‑Llama‑3.3‑70B model with a batch size of 4 and a sequence length of 16K tokens constitutes a single configuration. A bundle, however, can contain multiple such configurations, either for the same model or for different models. SambaNova’s RDU technology enables several models and configurations to be loaded simultaneously in a single deployment. This allows you to switch instantly between models and between batch‑/sequence‑size profiles as needed. In contrast to traditional GPU systems, where deployments are typically single‑model and static, SambaStack supports multi‑model, multi‑configuration bundles. This approach delivers higher efficiency, greater flexibility, and increased throughput while preserving low latency. You can run the following command to discover available bundles in your cluster:
The table below lists the recommended bundle templates for the models currently available in SambaStack. Each entry pairs a model with its recommended deployment bundle.
If the bundles listed below do not satisfy your inference requirements, you can create custom bundles that combine any mix of models and configurations so long as they fit in DDR memory.

Suggested bundles per model

For each model, this section lists the suggested bundle for typical use and any alternative bundles that trade off context length, batch size, or modality support. See Bundle configurations below for the seq length/batch size details of each bundle.

Meta

Meta-Llama-3.3-70B-InstructSuggested: 70b-3dot3-ss-4-8-16-32-64-128kAlternatives: 70b-3dot3-ss-full-whisper, us-agentic-rag-1-1, e5-mistral-70b-64k-128k
Meta-Llama-3.1-8B-InstructSuggested: us-agentic-rag-1-1Alternatives: qwen3-32b-llama405b-s-m
Meta-Llama-3.1-405B-InstructSuggested: qwen3-32b-llama405b-s-m
Llama-4-Maverick-17B-128E-InstructSuggested: llama-4-medium-8-16-32-64-128kAlternatives: llama-4-medium-ss-16k-bs24, us-agentic-rag-1-1

MiniMax

MiniMax-M2.7PREVIEWSuggested: dyt-minimax-m2p7-32-64-192k-pcAlternatives: dyt-minimax-m2p7-32-160-192k, dyt-minimax-m2p7-32k-v2
MiniMax-M2.5Suggested: dyt-minimax-m2p5-32-160kAlternatives: dyt-minimax-m2p5-32k

Mistral AI

Mistral-Large-3-675B-Instruct-2512PREVIEWSuggested: mistral-large-3-fp8-8-16-32kAlternatives: mistral-large-3-fp8-8k

DeepSeek

DeepSeek-R1-0528Suggested:
  • deepseek-5in1-fp8-16-32k (higher interactivity)
  • deepseek-4in1-fp8-128k (higher context length)
Alternatives: deepseek-r1-v3-fp8-8k, deepseek-r1-v31-fp8-8k
DeepSeek-V3-0324Suggested:
  • deepseek-5in1-fp8-16-32k (higher interactivity)
  • deepseek-4in1-fp8-128k (higher context length)
Alternatives: deepseek-r1-v3-fp8-8k, deepseek-v3-v31-fp8-8k, deepseek-v3-v3termi-fp8-8k
DeepSeek-V3.1Suggested:
  • deepseek-5in1-fp8-16-32k (higher interactivity)
  • deepseek-4in1-fp8-128k (higher context length)
Alternatives: deepseek-r1-v31-fp8-8k, deepseek-v3-v31-fp8-8k
DeepSeek-V3.1-TerminusSuggested:
  • deepseek-5in1-fp8-16-32k (higher interactivity)
  • deepseek-4in1-fp8-128k (higher context length)
Alternatives: deepseek-v3-v3termi-fp8-8k
DeepSeek-V3.2Suggested:
  • deepseek-5in1-fp8-16-32k (higher interactivity)
  • deepseek-4in1-fp8-128k (higher context length)

OpenAI

gpt-oss-120bSuggested: us-agentic-rag-1-1Alternatives: cd-dyt-gpt-oss-120b-8-32-64-128k, gpt-gemma-whisper-mistral
gpt-oss-20bPREVIEWSuggested: dyt-gpt-oss-20b-32-64-128k
Whisper-Large-v3Suggested: qwen3-32b-whisper-e5-mistralAlternatives: 70b-3dot3-ss-full-whisper (known issue: does not load within default startup time), gpt-gemma-whisper-mistral

Google

gemma-3-27b-itPREVIEWSuggested: gemma3-27b-32-128k
gemma-3-12b-itPREVIEWSuggested: gemma3-v3Alternatives: gpt-gemma-whisper-mistral
gemma-4-31B-itPREVIEWSuggested:
  • cd-gemma-4-31b-32-128-256k (text only; higher context length and constrained decoding support, no image/video support)
  • gemma-4-31b-32-128k (image/video support, higher throughput)

Alibaba Cloud

Qwen3-235B-A22B-Instruct-2507Suggested: dyt-qwen3-235b-32-128k
Qwen3-32BSuggested: qwen3-32b-whisper-e5-mistralAlternatives: qwen3-32b-llama405b-s-m
Qwen3-TTS-TalkerPREVIEWSuggested: qwen3-tts-talker
Qwen3-TTS-VocoderPREVIEWSuggested: qwen3-tts-vocoder

Other

E5-Mistral-7B-InstructSuggested: us-agentic-rag-1-1Alternatives: e5-mistral-70b-64k-128k, qwen3-32b-whisper-e5-mistral, gpt-gemma-whisper-mistral

Bundle configurations

The table below lists the configuration details for each bundle template referenced above. † These bundles support sequence lengths from 8K–128K. If you require shorter context lengths (4K–16K) for gpt-oss-120b, contact your SambaNova representative.