> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sambanova.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Checkpoint conversion tool

The checkpoint conversion packages convert HuggingFace-format checkpoints into the SambaNova-compatible format required for deployment on SambaRack and SambaCloud systems.

Conversion ships as Python packages, one per model family, installed from your SambaStack artifact registry. Each package converts checkpoints for the model family it is named for. The packages replace the Docker-based conversion tool.

<Note>
  Checkpoint conversion is a substep of deploying custom checkpoints on SambaStack or SambaCloud. See the [Deploying custom checkpoints](/docs/en/v2.2.3/sambastack/service-administration/model-deployment/deploying-models-and-bundles/deploy-custom-checkpoints) page for the high-level workflow.
</Note>

## Prerequisites

### System requirements

| Requirement | Specification |
| - | - |
| **Memory** | At least 2.5x the size of your checkpoint's SafeTensor or .bin files. For example, a 10GB checkpoint requires a minimum of 25GB available memory. |
| **Storage** | Equal to the input checkpoint size for the output. For example, converting a 14GB checkpoint requires 14GB of free storage for the converted output (plus the original). |
| **Operating System** | macOS or Linux-based operating systems. Windows OS is **not supported**. |
| **Python** | 3.11 or later. |

### Estimated conversion times

| Checkpoint Size | Example Model | Estimated Time |
| - | - | - |
| \~8GB | Meta-Llama-3.1-8B-Instruct | \~5 minutes |
| \~140GB | Meta-Llama-3.3-70B-Instruct | \~1 hour |

<Tip>
  Run the conversion locally on your machine or workspace that has access to your checkpoint's storage mount. A cloud compute instance can also be used, but note that data transfer of checkpoints can take a long time.
</Tip>

### Required software

* **Python** 3.11 or later, with `pip`
* **Google Cloud CLI** - [Installation guide](https://cloud.google.com/sdk/docs/install)

### Required access

* The registry URL for your SambaStack artifact registry, provided by your SambaNova account team or SambaNova support. You cannot install a conversion package without it
* Read access to that registry
* Authentication credentials for Google Cloud (the same account used for your organization's SambaStack artifact registry, if configured; otherwise contact your SambaNova account team or support)

<Note>
  Your access to SambaNova-hosted registries and checkpoint storage is read-only. You pull conversion packages from the registry and write your converted checkpoint to storage you control. See [Deploying custom checkpoints](/docs/en/v2.2.3/sambastack/service-administration/model-deployment/deploying-models-and-bundles/deploy-custom-checkpoints).
</Note>

## Supported models and checkpoint formats

### Supported model architectures

Custom checkpoints are supported for decoder-only text generation models. To confirm support for a model family and see the current exceptions, see the [Supported models](/docs/en/v2.2.3/sambastack/service-administration/model-deployment/supported-models-and-bundles) page.

### Checkpoint format requirements

Checkpoints are accepted in the HuggingFace format. The tensors should be in the safetensors format and the checkpoint directory should contain the same relevant config files as the base model for the custom checkpoint.

For example, if the custom checkpoint is a finetuned variant of [meta-llama/Llama-3.3-70B-Instruct](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct/tree/main), the checkpoint directory should contain files similar to the following:

| File | Description |
| - | - |
| `config.json` | Defines the model architecture configuration, including hidden size, number of layers, attention heads, and other structural parameters. |
| `generation_config.json` | Stores default text generation parameters such as sampling strategy, temperature, and maximum tokens. |
| `model-00001-of-00030.safetensors` ... `model-00030-of-00030.safetensors` | Sharded model weight files containing the trained tensor parameters of the model. Large models split weights across multiple files. |
| `model.safetensors.index.json` | Index file mapping model weight tensors to the corresponding safetensor shard files. |
| `special_tokens_map.json` | Defines special tokens used by the tokenizer such as BOS, EOS, and padding tokens. |
| `tokenizer.json` | Contains the tokenizer vocabulary and tokenization rules used to convert text into model input tokens. |
| `tokenizer_config.json` | Configuration metadata for the tokenizer, including tokenization behavior and settings. |

### Checkpoint compatibility

Given that a checkpoint is fine-tuned or derived from one of the supported models for your platform, checkpoints are compatible when their computational graph **has not been modified from the original checkpoint** (i.e., tensor weights and shapes).

**Aspects that must remain unchanged:**

* Number of attention heads
* Rope type (rope theta)
* Model vocabulary size
* Optimizer type
* Static architectural attributes in `config.json` such as: `head_dim`, `hidden_act`, `intermediate_size`, `attention_bias`, `attention_dropout`, `vocab_size`

**Aspects that can be modified:**

* Model weights or model weight tensor values
* Tokenizer and vocabulary, as long as the vocabulary size stays exactly the same as the original model checkpoint. This is useful for multilingual use cases.

<Tip>
  It can be helpful to think about this in terms of a static graph. Aspects of a model that are typically static in engines such as TensorRT-LLM are also static for custom checkpoints.
</Tip>

### Practical compatibility examples

Take the base model [meta-llama/Llama-3.3-70B-Instruct](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct) (a base model supported by SambaNova). The following checkpoints use the same computational graph as the original 70B model and **can** be converted and deployed on SambaNova platforms:

* [deepseek-ai/DeepSeek-R1-Distill-Llama-70B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B)
* [tokyotech-llm/Llama-3.3-Swallow-70B-Instruct-v0.4](https://huggingface.co/tokyotech-llm/Llama-3.3-Swallow-70B-Instruct-v0.4)

These checkpoints have undergone updates to their model weights, which have been adjusted and refined to improve performance or adapt to specific tasks or datasets.

## Choose your conversion package

Packages are named for the model family, not the individual model. Find the base model your checkpoint derives from.

| Base model | Package | Command |
| - | - | - |
| Meta-Llama-3.1 and 3.3 (all sizes), Llama-3.3-Swallow, DeepSeek-R1-Distill-Llama | `sn-conversion-llama` | `sn-convert-llama` |
| DeepSeek-V3 and DeepSeek-R1 (all variants) | `sn-conversion-deepseek-v3` | `sn-convert-deepseek-v3` |
| MiniMax-M2.5, MiniMax-M2.7 | `sn-conversion-minimax-m2` | `sn-convert-minimax-m2` |
| Mistral-Large-3 | `sn-conversion-mistral-large-3` | `sn-convert-mistral-large-3` |
| gpt-oss-120b, gpt-oss-20b | `sn-conversion-gpt-oss` | `sn-convert-gpt-oss` |
| Qwen3-235B-A22B | `sn-conversion-qwen3-moe` | `sn-convert-qwen3-moe` |

<Warning>
  Converting with the wrong package can produce a checkpoint that loads without error and returns incorrect output. Match the package to the base model your checkpoint derives from.
</Warning>

<Note>
  This table is the complete list of supported conversion packages. The registry holds other `sn-conversion-*` packages that are not supported for custom checkpoints and are not listed here. Use only the packages above.
</Note>

## Install the conversion package

<Steps>
  <Step stepNumber={1} titleSize="h3" title="Install and authenticate Google Cloud CLI">
    Install Google Cloud CLI in your conversion environment, following the official [Google Cloud CLI Installation](https://cloud.google.com/sdk/docs/install) guide, then authenticate:

    ```bash theme={} theme={}
    gcloud auth login
    ```

    <Note>
      Use the Google account associated with your organization's SambaStack artifact registry access.
    </Note>
  </Step>

  <Step stepNumber={2} titleSize="h3" title="Install the package">
    Install the authentication backend that lets `pip` read from the registry, then install the package for your model family:

    ```bash theme={} theme={}
    pip install keyring keyrings.google-artifactregistry-auth
    pip install --extra-index-url <REGISTRY_URL> sn-conversion-llama
    ```

    Replace `sn-conversion-llama` with the package for your model family from [Choose your conversion package](#choose-your-conversion-package). The examples on this page all use `sn-conversion-llama` and its `sn-convert-llama` command.

    <Note>
      Replace `<REGISTRY_URL>` with the registry URL provided by your SambaNova account team or SambaNova support. It takes the form `https://<REGION>-python.pkg.dev/<PROJECT>/<REPOSITORY>/simple/`. The registry holds only the conversion packages, so `--extra-index-url` keeps PyPI available for their dependencies.
    </Note>

    The shared conversion engine installs automatically as a dependency. Installing into a virtual environment is recommended.
  </Step>

  <Step stepNumber={3} titleSize="h3" title="Verify the installation">
    ```bash theme={} theme={}
    sn-convert-llama --help
    ```

    <Check>
      Setup is complete. To update to a newer version, run `pip install --upgrade` with your package name.
    </Check>
  </Step>
</Steps>

## Convert the checkpoint

### Command

Print the conversion plan first:

```bash theme={} theme={}
sn-convert-llama \
    --source_path /path/to/checkpoint \
    --target_path /path/to/converted \
    --print-plan
```

Drop `--print-plan` to run the conversion.

### Parameters

| Flag | Description |
| - | - |
| `-s`, `--source_path` | Path to the directory holding your input checkpoint. |
| `-t`, `--target_path` | Path where the converted checkpoint is written. Needs free space equal to the source checkpoint. |
| `-f`, `--force` | Overwrite a non-empty target directory. |
| `--print-plan` | Print the operations the conversion will perform, then exit. |
| `--dry-run` | Run the plan without writing output. Not available for every package, see [Known issues](#known-issues). |
| `--set` | Override a conversion setting. See [Conversion settings](#conversion-settings). |
| `--strict-ops` | Fail if a requested operation has no effect. See [Verify the conversion](#verify-the-conversion). |

<Warning>
  Type the flags exactly as shown. `--source_path` and `--target_path` use underscores, while `--print-plan`, `--dry-run`, and `--strict-ops` use hyphens. They are not interchangeable: both `--source-path` and `--strict_ops` are rejected.
</Warning>

### Conversion settings

Each package declares its own settings, and the defaults are correct for a checkpoint that matches its base model. `--print-plan` shows the resolved settings for your run.

Override a setting with `--set`. Boolean settings take a `+` or `-` prefix; others take `key=value`:

```bash theme={} theme={}
sn-convert-llama -s /path/to/checkpoint -t /path/to/converted \
    --set shard_bytes=5
```

These settings are common to every package:

| Setting | Description |
| - | - |
| `shard_bytes` | Maximum size in GB of each output safetensors shard. Defaults to 5. |
| `cache_bytes` | Memory budget in GB for tensors held during conversion. Lower this if you hit an out-of-memory condition. |
| `allow_symlink` | Symlink unchanged shards instead of copying them. Off by default; see the warning below. |

<Warning>
  Leave `allow_symlink` at its default. If any tensor shape changes during conversion, symlinked shards serve the original tensors and the checkpoint fails to load at deployment.
</Warning>

An unrecognized `--set` key fails the run and lists the keys the package accepts.

## Verify the conversion

### Check the output

A successful conversion writes the converted safetensors shards, an updated `model.safetensors.index.json`, and `sn_checkpoint_conversion_metadata.json` to the target directory, and exits zero.

`sn_checkpoint_conversion_metadata.json` records what each operation did. Check it when a conversion succeeds but the output looks wrong.

### Catch operations that did nothing

Re-run with `--strict-ops` to fail the conversion if a requested operation had no effect:

```bash theme={} theme={}
sn-convert-llama -s /path/to/checkpoint -t /path/to/converted --strict-ops
```

An operation with no effect usually means the checkpoint does not hold the tensors the package expected, which points to a mismatch between your checkpoint and its base model.

### Check against the PEF

Each PEF ships a `<pef_stem>_coe_meta.json` manifest alongside it, listing every tensor the compiled model expects with its shape and dtype.

To get the manifest, find the PEF's storage path in its [PEF resource](/docs/en/v2.2.3/sambastack/service-administration/model-deployment/custom-resources/pef) (`spec.versions.<version>.source`). The manifest is in the same folder, named after the PEF file without its `.pef` extension:

```bash theme={} theme={}
kubectl get pef <pef-name> -o yaml
gcloud storage cp gs://<path>/<pef_stem>_coe_meta.json .
```

Check your converted checkpoint against that manifest:

```bash theme={} theme={}
python -m sn_checkpoint_surgery.coe_meta_validation \
    /path/to/<pef_stem>_coe_meta.json \
    /path/to/converted
```

A missing tensor or a shape disagreement exits non-zero and names the tensor. Dtype differences are reported but do not fail the check, because the runtime converts dtypes at load.

<Note>
  This check compares tensor names, shapes, and dtypes. It confirms the converted checkpoint fits the model the PEF was built for; it does not verify that the conversion produced numerically correct values.
</Note>

### After conversion

Conversion rewrites the safetensors files but does not update `config.json`. If deployment reports a dimension or quantization mismatch, check `config.json` against the converted tensors.

## Troubleshooting

| Problem | Cause | Solution |
| - | - | - |
| Process exits with no error and no success message | Out-of-memory condition | Increase available memory to at least 2.5x the checkpoint size, or lower `cache_bytes` |
| `--set` key rejected | The setting does not exist for this package | Run `--print-plan` to see the settings the package accepts |
| Missing files reported at start | Incomplete checkpoint directory | Confirm `config.json`, all safetensors shards, and the tokenizer files are present |
| Checkpoint loads at deployment but output is wrong | Converted with the wrong package, or `allow_symlink` was overridden | Reconvert with the package matching your base model, at default settings |
| `sn-convert-<model>: command not found` | Package not installed in the active environment | Reinstall in your active virtual environment |

Errors carry a stable code, such as `CKPT-CNV-012`. Include the code when you contact SambaNova support.

## Known issues

**`--dry-run` fails for `sn-conversion-gpt-oss`, `sn-conversion-deepseek-v3`, and `sn-conversion-qwen3-moe`**

Conversion itself works for these packages. Only the `--dry-run` flag is affected, so leave it off and run the conversion. Use `--print-plan` to review the operations beforehand.

## Next steps

After successfully converting your checkpoint:

1. **Upload** the converted checkpoint to your GCS bucket or NFS mount
2. **Reference** the checkpoint from a Model resource or a checkpoint override
3. **Deploy** the checkpoint with a compatible ModelProfile

See [Deploying custom checkpoints](/docs/en/v2.2.3/sambastack/service-administration/model-deployment/deploying-models-and-bundles/deploy-custom-checkpoints) for the complete workflow.
