loupe

Model support

Which model families work with which part of loupe, and how each claim was checked.

loupe loads the unmodified Hugging Face model, so a family works when loupe can find its decoder blocks, final norm, attention and output projection. Real checkpoints are not run in CI: nothing downloads there.

MarkMeaning
testedCovered by the test suite or a --tiny example, on a tiny model offline
smokeRun once on a randomly initialised two-layer config of the architecture, not in CI
expectedThe code path is generic and should work; not run
noDoes not work today

Causal language models

FamilyInterploupe/ providerTool callsLoRA on TRLSAE features
Qwen2, Qwen2.5testedtestedtestedtestedtested
Qwen3smokeexpectedexpectedexpectedexpected
Llama 2, 3smokeexpectedexpectedexpectedexpected
Mistralsmokeexpectedexpectedexpectedexpected
Gemma 2, Gemma 3 1B (text only)smokeexpectedexpectedexpectedexpected
Phi-3smokeexpectedexpectedset targetsexpected
GPT-2smokechat templatechat templateset targetsexpected
GPT-NeoX, Pythiasmokechat templatechat templateset targetsexpected
Gemma 3 4B and up (vision)nono editsexpectedexpectedno
  • Interp is steer and ablate during generation, next-token logits, residual reads, logit lens, attention patterns, residual, head and attribution patching, and linear probes, all through nnsight. The smoke check ran each of these per family.
  • loupe/ provider renders every conversation with the tokenizer's chat template. GPT-2 and Pythia checkpoints ship without one: set tokenizer.chat_template or use an instruct fine-tune. The examples' chat() helper needs one too.
  • Tool calls go through the chat template's tools argument and are parsed by Inspect's Hugging Face handler, picked by the config's model_type: Llama and Mistral have their own parsers, every other family the generic one (<tool_call>, fenced JSON, <function>, bare JSON). A template that ignores tools leaves the model unaware of them. Tested with the Qwen format on a scripted reply.
  • LoRA on TRL targets q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj and down_proj by default. Phi-3 fuses them (qkv_proj, gate_up_proj), so the default adapts only o_proj and down_proj; GPT-2 (c_attn, c_proj, c_fc) and GPT-NeoX (query_key_value, dense) fail until lora.target_modules is set. SFT, DPO, GRPO and soft prompts are tested on the tiny Qwen2.
  • Unsloth replaces TRL's loading with 4-bit QLoRA and adds GGUF export, on a CUDA GPU with unsloth installed; backend: auto picks it then. Expected for the families Unsloth supports; not run, since CI has no GPU. Soft prompts and masked diffusion stay on TRL.
  • SAE features take any SAELens SAE on a residual hook (blocks.L.hook_resid_pre or hook_resid_post), read with nnsight on the unmodified model. Tested with a random SAE on the tiny Qwen2; a pretrained SAE exists only for some models, such as Gemma Scope for Gemma 2 and gpt2-small-res-jb for GPT-2.
  • Attention kernels: -M attn=sdpa (or eager, flash_attention_2, flex_attention, any name registered with transformers' AttentionInterface, or file.py:function to register one) picks the kernel, for the provider, load(attn=...) and loupe serve --attn. Only eager returns attention weights; attention views switch to it for their trace, head edits work under any kernel.
  • Gemma 3 4B and up load as a vision model whose decoder blocks sit at a path loupe does not look up yet, so it generates through the provider but takes no edit or analysis; the error names the paths it tried.

Masked diffusion models

ModelSampler and providerMasked diffusion SFT (LoRA)Adapter bank and phasesInterp
Any Hugging Face masked LM (BERT)testedtestedtestedno
LLaDAexpectedexpectedexpectedno
Dreamexpectedexpectedexpectedno

Served with -M diffusion='{"length": 64, "steps": 64}'. LLaDA and Dream run their own code from the Hub, so a Hub id needs -M revision=<commit>, the same arg that pins any Hub model. Steer, ablate and inject act on causal models only; the provider refuses them with diffusion. Training is TRL only.

Adding a family

A family whose blocks live elsewhere needs one path in BLOCK_PATHS, NORM_PATHS, ATTN_PATHS or OUT_PATHS in src/loupe/models/load.py.

On this page