Model support
Which model families work with which part of loupe, and how each claim was checked.
loupe loads the unmodified Hugging Face model, so a family works when loupe can find its decoder blocks, final norm, attention and output projection. Real checkpoints are not run in CI: nothing downloads there.
| Mark | Meaning |
|---|---|
| tested | Covered by the test suite or a --tiny example, on a tiny model offline |
| smoke | Run once on a randomly initialised two-layer config of the architecture, not in CI |
| expected | The code path is generic and should work; not run |
| no | Does not work today |
Causal language models
| Family | Interp | loupe/ provider | Tool calls | LoRA on TRL | SAE features |
|---|---|---|---|---|---|
| Qwen2, Qwen2.5 | tested | tested | tested | tested | tested |
| Qwen3 | smoke | expected | expected | expected | expected |
| Llama 2, 3 | smoke | expected | expected | expected | expected |
| Mistral | smoke | expected | expected | expected | expected |
| Gemma 2, Gemma 3 1B (text only) | smoke | expected | expected | expected | expected |
| Phi-3 | smoke | expected | expected | set targets | expected |
| GPT-2 | smoke | chat template | chat template | set targets | expected |
| GPT-NeoX, Pythia | smoke | chat template | chat template | set targets | expected |
| Gemma 3 4B and up (vision) | no | no edits | expected | expected | no |
- Interp is steer and ablate during generation, next-token logits, residual reads, logit lens, attention patterns, residual, head and attribution patching, and linear probes, all through nnsight. The smoke check ran each of these per family.
loupe/provider renders every conversation with the tokenizer's chat template. GPT-2 and Pythia checkpoints ship without one: settokenizer.chat_templateor use an instruct fine-tune. The examples'chat()helper needs one too.- Tool calls go through the chat template's
toolsargument and are parsed by Inspect's Hugging Face handler, picked by the config'smodel_type: Llama and Mistral have their own parsers, every other family the generic one (<tool_call>, fenced JSON,<function>, bare JSON). A template that ignorestoolsleaves the model unaware of them. Tested with the Qwen format on a scripted reply. - LoRA on TRL targets
q_proj,k_proj,v_proj,o_proj,gate_proj,up_projanddown_projby default. Phi-3 fuses them (qkv_proj,gate_up_proj), so the default adapts onlyo_projanddown_proj; GPT-2 (c_attn,c_proj,c_fc) and GPT-NeoX (query_key_value,dense) fail untillora.target_modulesis set. SFT, DPO, GRPO and soft prompts are tested on the tiny Qwen2. - Unsloth replaces TRL's loading with 4-bit QLoRA and adds GGUF export, on a CUDA GPU with
unslothinstalled;backend: autopicks it then. Expected for the families Unsloth supports; not run, since CI has no GPU. Soft prompts and masked diffusion stay on TRL. - SAE features take any SAELens SAE on a residual hook (
blocks.L.hook_resid_preorhook_resid_post), read with nnsight on the unmodified model. Tested with a random SAE on the tiny Qwen2; a pretrained SAE exists only for some models, such as Gemma Scope for Gemma 2 andgpt2-small-res-jbfor GPT-2. - Attention kernels:
-M attn=sdpa(oreager,flash_attention_2,flex_attention, any name registered with transformers'AttentionInterface, orfile.py:functionto register one) picks the kernel, for the provider,load(attn=...)andloupe serve --attn. Only eager returns attention weights; attention views switch to it for their trace, head edits work under any kernel. - Gemma 3 4B and up load as a vision model whose decoder blocks sit at a path loupe does not look up yet, so it generates through the provider but takes no edit or analysis; the error names the paths it tried.
Masked diffusion models
| Model | Sampler and provider | Masked diffusion SFT (LoRA) | Adapter bank and phases | Interp |
|---|---|---|---|---|
| Any Hugging Face masked LM (BERT) | tested | tested | tested | no |
| LLaDA | expected | expected | expected | no |
| Dream | expected | expected | expected | no |
Served with -M diffusion='{"length": 64, "steps": 64}'. LLaDA and Dream run their own code from
the Hub, so a Hub id needs -M revision=<commit>, the same arg that pins any Hub model. Steer, ablate and inject act on causal models
only; the provider refuses them with diffusion. Training is TRL only.
Adding a family
A family whose blocks live elsewhere needs one path in BLOCK_PATHS, NORM_PATHS, ATTN_PATHS or
OUT_PATHS in src/loupe/models/load.py.