loupe

Research domains

What loupe can test, organised by question, not by project.

DomainCapabilities
Mechanismsdirections, activation and attribution patching, logit lens, probes, attention, SAE features, circuits
Evaluation sciencegrids over seeds with paired intervals, moved and held verdicts, black-box endpoints, contamination checks
Conditioningprompt text, few-shot, steering vectors and soft prompts reaching one behaviour target
Adaptersa bank of LoRA adapters live in any subset, PEFT merges, per-site overlap
Decoding and model familiesper-step hooks, adapters switched by generation phase, masked diffusion models
Retrieval and groundingBM25, dense and hybrid search, reranking; recall, F1, exact match, NLI faithfulness
Retrieval inside the modelpassages injected at a layer, retrieval during decoding, spliced KV divergence
Small models and datateacher collection, hashed splits, n-gram contamination, SFT, a small classifier

Examples in the repository

ExperimentQuestion
sycophancy-pushbackIs caving to unsupported pushback a direction, and does removing it keep correctness?
refusal-directionIs refusal mediated by one direction (Arditi et al., 2024)?
refusal-finetuningDoes fine-tuning move that direction?
gsm8k-grpo, tool-rlGRPO with one check as reward and scorer; tool environments
agent-sandboxAn agent with bash in Docker, under interventions

On this page