Skip to main content
SambaStack supports two deployment configurations for supported models: high-interactivity and high-throughput. Both use the same model weights and the same API – they differ only in how the system handles requests. Use high-interactivity for low-latency, user-facing applications. Use high-throughput for batch or high-concurrency workloads where aggregate output matters more than per-user latency. The rest of this page covers the trade-offs, the model profiles and PEF configurations for each, and how to deploy them.
High-throughput and high-interactivity configurations require dedicated systems. Models deployed in either configuration cannot be bundled with other models. If you are unfamiliar with models, profiles, and bundles, see Deploying models and bundles.

Deployment configurations

Both configurations use the same model name in API calls. The same request works against either configuration:
The configuration controls request handling on the server side; no client-side changes are required.

When to use each configuration

High-throughput

Use the high-throughput configuration when:
  • You are serving many concurrent users and aggregate throughput matters more than per-user latency
  • Your workload is asynchronous or batch-oriented (for example, document processing or offline inference pipelines)
  • End-to-end latency per request is not a constraint

High-interactivity

Use the high-interactivity configuration when:
  • You are building real-time, user-facing applications
  • Per-user time-to-first-token and tokens-per-second are the primary metrics
  • Your deployment has fewer nodes, or your users have tight latency budgets

Supported models

Both configurations are available for the following models:
  • DeepSeek-R1
  • DeepSeek-V3-0324
  • DeepSeek-V3.1
  • DeepSeek-V3.1-Terminus
  • DeepSeek-V3.2

Architecture

The high-throughput configuration uses continuous batching, separating the prefill and decode phases into a dedicated pipeline. Two modes are available:
  • Aggregated (ACB): Prefill and decode run collocated on the same nodes.
  • Disaggregated (DCB): Prefill and decode run on separate dedicated nodes, so each phase can be sized independently. The recommended node split is more prefill nodes than decode nodes – for example, three prefill nodes and one decode node.
DCB has not yet been internally validated on SambaStack. Use ACB for SambaStack deployments until DCB validation is published.
  • Prefill nodes process the input prompt
  • Decode nodes generate output tokens
The high-throughput configuration requires a minimum of 4 nodes in disaggregated mode. For single-node or small deployments, use the high-interactivity configuration instead.

Requirements and limitations

PEF configurations

Each configuration is delivered as a ModelProfile. You select a profile rather than an individual PEF, and the profile determines which PEF, sequence length, and batch size are used. The PEF tables below document which combinations each configuration provides. See Deploying models and bundles for the full deployment procedure.
Custom Resource (CR): A Kubernetes extension object. Model, ModelProfile, and Pef resources define the models, runtime configurations, and compiled executables available in the cluster.
A ModelProfile lists its PEFs in spec.pefs as <pef-name>[:<version>], for example deepseek-ss8192-bs1:1. You do not normally edit this list; use it to confirm which PEFs a profile provides.

High-throughput PEFs

Picking a high-throughput PEF:
  • Choose the sequence length (ss) that fits your longest prompt plus expected output tokens.
  • Higher batch sizes serve more concurrent decode requests per node but require more RDU memory. The table lists the supported combinations.

High-interactivity PEFs

Picking a high-interactivity PEF:
  • Match the sequence length to your prompt plus expected output budget.
  • bs1 minimizes per-user latency. bs4 trades a small latency increase for higher per-node throughput when you have multiple concurrent users.

Deploy a configuration

No prebuilt bundles ship for these configurations. Select the ModelProfile that corresponds to the configuration you want, then deploy it. Continuous batching is a property of the compiled PEF, surfaced in the profile’s spec.features, so you do not enable it yourself. For DeepSeek, the two profiles are: For any other model, identify the profile by checking which ones report continuous_batching in spec.features:
A profile whose features includes continuous_batching provides the high-throughput configuration. A profile without it provides the high-interactivity configuration. Cross-check the profile’s spec.pefs against the PEF tables above to confirm the sequence lengths and batch sizes it serves. Pair the model with that profile in a ModelBundle, or inline it in a ModelDeployment for a single model:
Then configure the ModelDeployment replica groups for the appropriate mode. Aggregated mode (ACB):
Disaggregated mode (DCB):

Verify your deployment

After deploying, confirm the configuration is active:
In the output, look for continuous_batching.mode set to aggregate (ACB) or disaggregate (DCB), and confirm the replica counts under prefill and decode match what you configured.

Switch between configurations

To switch between high-throughput and high-interactivity, redeploy with the profile for the target configuration. When switching to high-interactivity, reference a profile whose features does not include continuous_batching, and remove the continuous_batching block from the ModelDeployment replica group. The model name in API calls does not change.

Monitor your deployment

The SambaStack logging system emits per-request metrics relevant to these deployments: See Logs for the full list of available metrics and example queries.