DeepSeek-R1-0528 /
DeepSeek-V3.1 | deepseek-r1-v31-fp8-8k | - Medium context length with low batch size
| ViewModels: DeepSeek-R1-0528DeepSeek-V3.1
|
DeepSeek-R1-0528 /
DeepSeek-V3-0324 | deepseek-r1-v3-fp8-8k | - Medium context length with low batch size
| ViewModels: DeepSeek-R1-0528DeepSeek-V3-0324
|
DeepSeek-V3-0324 /
DeepSeek-V3.1 | deepseek-v3-v31-fp8-8k | - Medium context length with low batch size
| ViewModels: DeepSeek-V3-0324DeepSeek-V3.1
|
DeepSeek-V3-0324 /
DeepSeek-V3.1-Terminus | deepseek-v3-v3termi-fp8-8k | - Medium context length with low batch size
| ViewModels: DeepSeek-V3-0324DeepSeek-V3.1-Terminus
|
DeepSeek-R1-0528 /
DeepSeek-V3-0324 /
DeepSeek-V3.1 /
DeepSeek-V3.1-Terminus | deepseek-4in1-fp8-128k | - Large context length with single batch size
- Four DeepSeek models in one bundle
| ViewModels: DeepSeek-R1-0528DeepSeek-V3-0324DeepSeek-V3.1DeepSeek-V3.1-Terminus
|
E5-Mistral-7B-Instruct /
Meta-Llama-3.1-8B-Instruct /
Llama-4-Maverick-17B-128E-Instruct /
Meta-Llama-3.3-70B-Instruct /
gpt-oss-120b | us-agentic-rag-1-1 | - Small to medium context length with varied batch size
- Speculative decoding supported for
Meta-Llama-3.3-70B
| Viewgpt-oss-120b- Seq Length: 32K, BS: 4
- Seq Length: 64K, BS: 2
- Seq Length: 128K, BS: 2
Llama-4-Maverick-17B-128E-Instruct- Seq Length: 8K, BS: 1
- Seq Length: 16K, BS: 1
Meta-Llama-3.3-70B (Target)/ Meta-Llama-3.2-1B (Draft)- Seq Length: 4K, BS: 1, 4, 8, 16, 32
- Seq Length: 8K, BS: 1, 4, 8
- Seq Length: 16K, BS: 1, 4
- Seq Length: 32K, BS: 1, 4
- Seq Length: 64K, BS: 1
- Seq Length: 128K, BS: 1
Meta-Llama-3.1-8B-Instruct- Seq Length: 4K, BS: 1, 4, 16, 32
- Seq Length: 8K, BS: 1, 4, 16, 32
- Seq Length: 16K, BS: 1, 4, 8
E5-Mistral-7B-Instruct- Seq Length: 4K, BS: 1, 4, 8, 16, 32
|
E5-Mistral-7B-Instruct /
Meta-Llama-3.3-70B | e5-mistral-70b-64k-128k | - Large context length with low batch size
- Speculative decoding supported for
Meta-Llama-3.3-70B
| ViewModels: E5-Mistral-7B-Instruct- Seq Length: 4K, BS: 1, 4, 8, 16, 32
Meta-Llama-3.3-70B (Target)/ Meta-Llama-3.2-1B (Draft)- Seq Length: 64K, BS: 1
- Seq Length: 128K, BS: 1
|
gemma-3-27b-it | gemma3-27b-32-128k | Homogeneous bundle containing gemma-3-27b-it configurations. | Viewgemma-3-27b-it- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 128K, BS: 2, 4, 6, 8
|
gemma-3-12b-it | gemma3-v3 | - Homogeneous bundle containing
gemma-3-12b-it configurations. - Large context length with medium batch size
| Viewgemma-3-12b-it- Seq Length: 128K, BS: 2, 4, 6, 8
|
gemma-4-31B-it | gemma-4-31b-32-128k | Homogeneous bundle with constrained decoding for gemma-4-31B-it. Supports text, image, and video input. | Viewgemma-4-31B-it- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 128K, BS: 2, 4, 6, 8
|
gemma-4-31B-it | gemma-4-31b-32-128-256k | - Homogeneous bundle with constrained decoding for
gemma-4-31B-it. Adds context support up to 256K. Supports text, image, and video input at all sequence lengths.
| Viewgemma-4-31B-it- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 128K, BS: 2, 4, 6, 8
- Seq Length: 256K, BS: 2
|
gemma-4-31B-itPREVIEW | gemma-4-31b-mtp-cd-8-32-64-128k | - Homogeneous bundle with multi-token prediction and constrained decoding for
gemma-4-31B-it. Up to 128K context; text, image, and video input at all sequence lengths. - Multi-token prediction raises decode throughput; the model predicts several tokens per step and verifies them in the same pass.
| Viewgemma-4-31B-it- Seq Length: 8K, BS: 2, 4, 6, 8
- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 64K, BS: 2, 4, 6, 8
- Seq Length: 128K, BS: 2, 4, 6, 8
|
gpt-oss-120b | cd-dyt-gpt-oss-120b-8-32-64-128k † | - Homogeneous bundle with constrained decoding for
gpt-oss-120b. - Covers 8K through 128K context.
- Includes structured output (logit-masking) support.
| Viewcd-dyt-gpt-oss-120b-8-32-64-128k- Seq Length: 8K, BS: 2, 4, 6, 8
- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 64K, BS: 2, 4
- Seq Length: 128K, BS: 2
|
gpt-oss-20b | dyt-gpt-oss-20b-32-64-128k † | Homogeneous bundle with constrained decoding for gpt-oss-20b. | Viewdyt-gpt-oss-20b-32-64-128k- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 64K, BS: 2, 4
- Seq Length: 128K, BS: 2
|
E5-Mistral-7B-Instruct /
Whisper-Large-v3 /
gemma-3-12b-it /
gpt-oss-120b | gpt-gemma-whisper-mistral | - Small to medium context length combining embeddings, transcription, image understanding, and a tool-calling agent
| ViewModels: E5-Mistral-7B-InstructWhisper-Large-v3gemma-3-12b-it- Seq Length: 128K, BS: 2, 4, 6, 8
gpt-oss-120b- Seq Length: 8K, BS: 2, 4, 6, 8
|
Llama-4-Maverick-17B-128E-Instruct | llama-4-medium-8-16-32-64-128k | - Homogeneous bundles containing
Llama-4-Maverick-17B-128E-Instruct configurations. - Small to large context length with low batch
| ViewLlama-4-Maverick-17B-128E-Instruct- Seq Length: 8K, BS: 1
- Seq Length: 16K, BS: 1
- Seq Length: 32K, BS: 1
- Seq Length: 64K, BS: 1
- Seq Length: 128K, BS: 1
|
Llama-4-Maverick-17B-128E-Instruct | llama-4-medium-ss-16k-bs24 | Homogeneous bundle containing Llama-4-Maverick-17B-128E-Instruct configurations; medium context length with higher batch size. | ViewLlama-4-Maverick-17B-128E-Instruct- Seq Length: 16K, BS: 2, 4
|
Meta-Llama-3.3-70B-Instruct | 70b-3dot3-ss-4-8-16-32-64-128k | - Medium to large context length with low batch size
| ViewTarget Models: Meta-Llama-3.3-70B-Instruct- Seq Length: 4K, BS: 2, 4, 8, 16, 32
- Seq Length: 8K, BS: 2, 4, 8, 16, 32
- Seq Length: 16K, BS: 1, 2, 4
- Seq Length: 32K, BS: 1, 2, 4
- Seq Length: 64K, BS: 1, 2, 4
- Seq Length: 128K, BS: 1
Draft Models: Meta-Llama-3.2-1B-Instruct- Seq Length: 4K, BS: 2, 4, 8, 16, 32
- Seq Length: 8K, BS: 2, 4, 8, 16, 32
- Seq Length: 16K, BS: 1, 2, 4
- Seq Length: 32K, BS: 1, 2, 4
- Seq Length: 64K, BS: 1, 2, 4; private: true
- Seq Length: 128K, BS: 1; private: true
|
Meta-Llama-3.3-70B-Instruct /
Whisper-Large-v3 | 70b-3dot3-ss-full-whisper | - Full context-length range with low to medium batch size, plus Whisper transcription
- Speculative decoding supported for
Meta-Llama-3.3-70B-Instruct
| ViewTarget Models: Meta-Llama-3.3-70B-Instruct- Seq Length: 4K, BS: 2, 4, 8, 16, 32
- Seq Length: 8K, BS: 2, 4, 8, 16, 32
- Seq Length: 16K, BS: 1, 2, 4
- Seq Length: 32K, BS: 1, 2, 4
- Seq Length: 64K, BS: 1, 2, 4
- Seq Length: 128K, BS: 1
Whisper-Large-v3
Draft Models: Meta-Llama-3.2-1B-Instruct- Seq Length: 4K, BS: 2, 4, 8, 16, 32
- Seq Length: 8K, BS: 2, 4, 8, 16, 32
- Seq Length: 16K, BS: 1, 2, 4
- Seq Length: 32K, BS: 1, 2, 4
- Seq Length: 64K, BS: 1, 2, 4
- Seq Length: 128K, BS: 1
|
MiniMax-M2.5 | - dyt-minimax-m2p5-32k
- dyt-minimax-m2p5-32-160k
| - Homogeneous bundles containing
MiniMax-M2.5 configurations. - dyt-minimax-m2p5-32k is better for medium sequence lengths and high batching.
- dyt-minimax-m2p5-32-160k is better for higher sequence lengths and low batching
| View- dyt-minimax-m2p5-32k
- Seq Length: 4K-32K, BS: 2, 4, 6, 8
- dyt-minimax-m2p5-32-160k
- Seq Length: 32K, BS: 2
- Seq Length: 160K, BS: 2
|
MiniMax-M2.7 | dyt-minimax-m2p7-32k-v2 | Homogeneous bundle containing MiniMax-M2.7 configurations; medium context length with high batching. | ViewMiniMax-M2.7- Seq Length: 8K-32K, BS: 2, 4, 6, 8
|
MiniMax-M2.7 | dyt-minimax-m2p7-32-160-192k | Homogeneous bundle containing MiniMax-M2.7 configurations; better for higher sequence lengths and low batching. | ViewMiniMax-M2.7- Seq Length: 8K-32K, BS: 2, 4, 6, 8
- Seq Length: 160K-192K, BS: 2
|
MiniMax-M2.7 | dyt-minimax-m2p7-32-64-160-192k-pc | Homogeneous bundle with prompt caching for MiniMax-M2.7. | ViewMiniMax-M2.7- Seq Length: 8K-32K, BS: 2, 4, 6, 8
- Seq Length: 64K, BS: 2, 4
- Seq Length: 160K-192K, BS: 2
|
MiniMax-M2.7 | dyt-minimax-m2p7-32-64-160-192k | - Homogeneous bundle containing
MiniMax-M2.7 configurations; same context range as dyt-minimax-m2p7-32-64-160-192k-pc without prompt caching.
| ViewMiniMax-M2.7- Seq Length: 8K-32K, BS: 2, 4, 6, 8
- Seq Length: 64K, BS: 2
- Seq Length: 160K-192K, BS: 2
|
MiniMax-M2.7 | dyt-minimax-m2p7-32k-pc | Homogeneous bundle with prompt caching for MiniMax-M2.7; medium context length with high batching. | ViewMiniMax-M2.7- Seq Length: 8K-32K, BS: 2, 4, 6, 8
|
MiniMax-M3PREVIEW | minimax-m3-32k | Homogeneous bundle containing MiniMax-M3 configurations; medium context length. Supports text, image, and video input. | ViewMiniMax-M3- Seq Length: 8K-32K, BS: 1, 2, 4
|
MiniMax-M3PREVIEW | minimax-m3-512k | - Homogeneous bundle containing
MiniMax-M3 configurations. - Up to 512K context; text, image, and video input at all sequence lengths.
| ViewMiniMax-M3- Seq Length: 8K-32K, BS: 1, 2, 4
- Seq Length: 64K, BS: 1, 2, 4
- Seq Length: 128K, BS: 1, 2, 4
- Seq Length: 256K, BS: 1, 2, 4
- Seq Length: 512K, BS: 1
|
MiniMax-M3PREVIEW | minimax-m3-32-64-128-256-512k-1m | - Homogeneous bundle containing
MiniMax-M3 configurations. - Up to 1M context; text, image, and video input at all sequence lengths.
| ViewMiniMax-M3- Seq Length: 8K-32K, BS: 1, 2, 4
- Seq Length: 64K, BS: 1, 2, 4
- Seq Length: 128K, BS: 1, 2, 4
- Seq Length: 256K, BS: 1, 2, 4
- Seq Length: 512K, BS: 1
- Seq Length: 1M, BS: 1
|
MiniMax-M3PREVIEW | minimax-m3-acb-prefix-caching | - Homogeneous bundle for
MiniMax-M3 with prompt caching, continuous batching, and constrained decoding. Up to 1M context. - Text input only. The other
MiniMax-M3 bundles serve text, image, and video.
| View |
Mistral-Large-3-675B-Instruct-2512PREVIEW | mistral-large-3-fp8-8k | - Homogeneous bundle containing
Mistral-Large-3-675B-Instruct-2512 configurations. - Preview model – text-only.
| ViewMistral-Large-3-675B-Instruct-2512
|
Mistral-Large-3-675B-Instruct-2512PREVIEW | mistral-large-3-fp8-8-16-32k | - Homogeneous bundle containing
Mistral-Large-3-675B-Instruct-2512 configurations. - Wider context range than
mistral-large-3-fp8-8k. Preview model – text-only.
| ViewMistral-Large-3-675B-Instruct-2512- Seq Length: 8K, BS: 1, 4
- Seq Length: 16K, BS: 1, 2
- Seq Length: 32K, BS: 1
|
Qwen3-235B-A22B-Instruct-2507 | dyt-qwen3-235b-32-128k | Homogeneous bundle containing Qwen3-235B-A22B-Instruct-2507 configurations. | ViewQwen3-235B-A22B-Instruct-2507- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 128K, BS: 2
|
Whisper-Large-v3 /
Qwen3-32B /
E5-Mistral-7B-Instruct | qwen3-32b-whisper-e5-mistral | - Small to medium context length with varied batch size
| ViewE5-Mistral-7B-Instruct- Seq Length: 4K, BS: 1, 4, 8, 16, 32
Qwen3-32B- Seq Length: 8K, BS: 1, 4
- Seq Length: 16K, BS: 1
- Seq Length: 32K, BS: 1, 2
Whisper-Large-v3
|
Qwen3-32B /
Meta-Llama-3.1-405B-Instruct | qwen3-32b-llama405b-s-m | - Small to medium context length
- Speculative decoding supported for
Meta-Llama-3.1-405B-Instruct
| ViewTarget Models: Meta-Llama-3.1-405B-Instruct- Seq Length: 4K, BS: 1, 2, 4
- Seq Length: 8K, BS: 1
- Seq Length: 16K, BS: 1
Draft Models: Meta-Llama-3.1-8B-InstructMeta-Llama-3.2-3B-Instruct- Seq Length: 4K, BS: 1, 2, 4
- Seq Length: 8K, BS: 1
Routable Models: Qwen3-32B- Seq Length: 8K, BS: 1, 4
- Seq Length: 16K, BS: 1
- Seq Length: 32K, BS: 1
|
Qwen3-TTS-TalkerPREVIEW | qwen3-tts-talker | Homogeneous bundle containing Qwen3-TTS-Talker configurations. | ViewQwen3-TTS-Talker- Seq Length: 4K, BS: 2, 4, 8, 16
|
Qwen3-TTS-VocoderPREVIEW | qwen3-tts-vocoder | Homogeneous bundle containing Qwen3-TTS-Vocoder configurations. | ViewQwen3-TTS-Vocoder- Codes length: 5, 8, 10, 16, BS: 16
|