Research domains
What loupe can test, organised by question, not by project.
| Domain | Capabilities |
|---|---|
| Mechanisms | directions, activation and attribution patching, logit lens, probes, attention, SAE features, circuits |
| Evaluation science | grids over seeds with paired intervals, moved and held verdicts, black-box endpoints, contamination checks |
| Conditioning | prompt text, few-shot, steering vectors and soft prompts reaching one behaviour target |
| Adapters | a bank of LoRA adapters live in any subset, PEFT merges, per-site overlap |
| Decoding and model families | per-step hooks, adapters switched by generation phase, masked diffusion models |
| Retrieval and grounding | BM25, dense and hybrid search, reranking; recall, F1, exact match, NLI faithfulness |
| Retrieval inside the model | passages injected at a layer, retrieval during decoding, spliced KV divergence |
| Small models and data | teacher collection, hashed splits, n-gram contamination, SFT, a small classifier |
Examples in the repository
| Experiment | Question |
|---|---|
sycophancy-pushback | Is caving to unsupported pushback a direction, and does removing it keep correctness? |
refusal-direction | Is refusal mediated by one direction (Arditi et al., 2024)? |
refusal-finetuning | Does fine-tuning move that direction? |
gsm8k-grpo, tool-rl | GRPO with one check as reward and scorer; tool environments |
agent-sandbox | An agent with bash in Docker, under interventions |