DeepSeek-R1-0528 /
DeepSeek-V3.1 | deepseek-r1-v31-fp8-16k | - Combination of:
DeepSeek-R1-0528 DeepSeek-V3.1
- Medium context length with low batch size
| ViewModels: DeepSeek-R1-0528DeepSeek-V3.1
|
DeepSeek-V3-0324 | deepseek-r1-v3-fp8-16k | - Combination of:
DeepSeek-R1-0528 DeepSeek-V3-0324
- Medium context length with low batch size
| ViewModels: DeepSeek-R1-0528DeepSeek-V3-0324
|
E5-Mistral-7B-Instruct /
Meta-Llama-3.1-8B-Instruct | us-agentic-rag-1-1 | - Combination of:
gpt-oss-120bLlama-4-Maverick-17B-128E-InstructMeta-Llama-3.1-8B-InstructMeta-Llama-3.3-70B (Target)Meta-Llama-3.2-1B (Draft)E5-Mistral-7B-Instruct
- Small to medium context length with varied batch size
- Speculative decoding supported for
Meta-Llama-3.3-70B
| Viewgpt-oss-120b- Seq Length: 32K, BS: 4
- Seq Length: 64K, BS: 2
- Seq Length: 128K, BS: 2
Llama-4-Maverick-17B-128E-Instruct- Seq Length: 8K, BS: 1
- Seq Length: 16K, BS: 1
Meta-Llama-3.3-70B (Target)/ Meta-Llama-3.2-1B (Draft)- Seq Length: 4K, BS: 1, 4, 8, 16, 32
- Seq Length: 8K, BS: 1, 4, 8
- Seq Length: 16K, BS: 1, 4
- Seq Length: 32K, BS: 1, 4
- Seq Length: 64K, BS: 1
- Seq Length: 128K, BS: 1
Meta-Llama-3.1-8B-Instruct- Seq Length: 4K, BS: 1, 4, 16, 32
- Seq Length: 8K, BS: 1, 4, 16, 32
- Seq Length: 16K, BS: 1, 4, 8
E5-Mistral-7B-Instruct- Seq Length: 4K, BS: 1, 4, 8, 16, 32
|
gemma-3-27b-it | gemma3-27b-32-128k | Homogeneous bundle containing gemma-3-27b-it configurations. | Viewgemma-3-27b-it- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 128K, BS: 2, 4, 6, 8
|
gemma-3-12b-it | gemma3-v3 | - Homogeneous bundle containing
gemma-3-12b-it configurations. - Large context length with medium batch size
| Viewgemma-3-12b-it- Seq Length: 128K, BS: 2, 4, 6, 8
|
gpt-oss-120b | cd-dyt-gpt-oss-120b-32-64-128k †cd-dyt-gpt-oss-120b-8-32-64-128k †
| - Homogeneous bundles with constrained decoding for
gpt-oss-120b. cd-dyt-gpt-oss-120b-8-32-64-128k adds 8K context support.
| Viewcd-dyt-gpt-oss-120b-32-64-128k- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 64K, BS: 2, 4
- Seq Length: 128K, BS: 2
cd-dyt-gpt-oss-120b-8-32-64-128k- Seq Length: 8K, BS: 2, 4, 6, 8
- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 64K, BS: 2, 4
- Seq Length: 128K, BS: 2
|
Llama-4-Maverick-17B-128E-Instruct | llama-4-medium-8-16-32-64-128k | - Homogeneous bundles containing
Llama-4-Maverick-17B-128E-Instruct configurations. - Small to large context length with low batch
| ViewLlama-4-Maverick-17B-128E-Instruct- Seq Length: 8K, BS: 1
- Seq Length: 16K, BS: 1
- Seq Length: 32K, BS: 1
- Seq Length: 64K, BS: 1
- Seq Length: 128K, BS: 1
|
Meta-Llama-3.3-70B-Instruct | 70b-3dot3-ss-4-8-16-32-64-128k | - Speculative decoding of:
Meta-Llama-3.3-70B (Target)Meta-Llama-3.2-1B (Draft)
- Medium to large context length with low batch size
| ViewTarget Models: Meta-Llama-3.3-70B-Instruct- Seq Length: 4K, BS: 2, 4, 8, 16, 32
- Seq Length: 8K, BS: 2, 4, 8, 16, 32
- Seq Length: 16K, BS: 1, 2, 4
- Seq Length: 32K, BS: 1, 2, 4
- Seq Length: 64K, BS: 1, 2, 4
- Seq Length: 128K, BS: 1
Draft Models: Meta-Llama-3.2-1B-Instruct- Seq Length: 4K, BS: 2, 4, 8, 16, 32
- Seq Length: 8K, BS: 2, 4, 8, 16, 32
- Seq Length: 16K, BS: 1, 2, 4
- Seq Length: 32K, BS: 1, 2, 4
- Seq Length: 64K, BS: 1, 2, 4; private: true
- Seq Length: 128K, BS: 1; private: true
|
MiniMax-M2.5 | - dyt-minimax-m2p5-32k
- dyt-minimax-m2p5-32-160k
| - Homogeneous bundles containing
MiniMax-M2.5 configurations. - dyt-minimax-m2p5-32k is better for medium sequence lengths and high batching.
- dyt-minimax-m2p5-32-160k is better for higher sequence lengths and low batching
| View- dyt-minimax-m2p5-32k
- Seq Length: 4K-32K, BS: 2, 4, 6, 8
- dyt-minimax-m2p5-32-160k
- Seq Length: 32K, BS: 2
- Seq Length: 160K, BS: 2
|
Qwen3-235B-A22B-Instruct-2507 | dyt-qwen3-235b-32-128k | Homogeneous bundle containing Qwen3-235B-A22B-Instruct-2507 configurations. | ViewQwen3-235B-A22B-Instruct-2507- Seq Length: 32K, BS: 2, 4, 6, 8
- Seq Length: 128K, BS: 2
|
Qwen3-235B | - qwen3-235b-8k
- qwen3-235b-16-32-64k
- qwen3-235b-128k
| - Homogeneous bundles containing
Qwen3-235B configurations. - qwen3-235b-8k: small context with higher batch size.
- qwen3-235b-16-32-64k: medium context lengths.
- qwen3-235b-128k: large context length.
| View- qwen3-235b-8k
- qwen3-235b-16-32-64k
- Seq Length: 16K, BS: 2
- Seq Length: 32K, BS: 2
- Seq Length: 64K, BS: 2
- qwen3-235b-128k
|
Whisper-Large-v3 /
Qwen3-32B | qwen3-32b-whisper-e5-mistral | - Combination of:
Qwen3-32BWhisper-Large-v3E5-Mistral-7B-Instruct
- Small to medium context length with varied batch size
| ViewE5-Mistral-7B-Instruct- Seq Length: 4K, BS: 1, 4, 8, 16, 32
Qwen3-32B- Seq Length: 8K, BS: 1, 4
- Seq Length: 16K, BS: 1
- Seq Length: 32K, BS: 1, 2
Whisper-Large-v3
|