I study how language models reason and respond to training, and put AI agents to work on hard, checkable problems: formal proofs, faster code, real bugs.
Write-ups, models and every repo: casella.dev
Each write-up lists its model, data, code and open issues.
- Verified once, reused three times: agents carried proofs of shipped OpenSSL code into new callers · sources
Machine-checked contracts for two unchanged Debian OpenSSL routines verified three experimental callers through their actual linkage. The held-out caller missed its three-hour cap, then completed in a recorded successor. - Agents turned a one-kernel Lean proof into a checker that certified 25 PyTorch kernels · code
One Lean theorem and a small checker certified 32-bit size arguments for 25torch.compilekernels, including 10 of 24 held out, with each proof bound to the live compilation. Two of nine tuned kernels got faster. - A Lean proof let agents cut a compiled PyTorch workload’s runtime by 26% · code · PR
Agents declared atorch.compilekernel’s size arguments 32-bit, proved in Lean when that is exact, and cut a changing-shape workload’s time by 26% on one L4. - 25 bugs in free-threaded CPython, found by agents
Agents confirmed 25 bugs in the no-GIL build, with no earlier report found for 18. TLAPS and Lean proofs of the locking model held; replays of real runs showed five places CPython leaves it. - I used formal verification and agents to make an algorithm 105× faster and prove it behaves the same
Seven rewrites of deliberately slow Dafny routines verified against a frozen spec; one ran 105× faster on a workload it never saw. - A 50-step RL update reduces categorical sampling bias · code
Fifty RL steps on one random-integer task moved nine untrained pick-one tasks toward uniform on Qwen3-30B-A3B-Instruct. - Hidden-state probes outperform self-reported confidence · code
A linear probe on Llama 3.1 8B’s hidden states ranks claim correctness better than the model’s stated confidence. - Panel-style reasoning trades accuracy for shorter completions and one RL run on sometimes-solvable problems · code
All nine: casella.dev/research.html
Machine-checked software. A model proposes; a proof assistant or the toolchain decides what ships.
-
HN, formally · live site — The Hacker News front page and its 30 threads, rendered by a Lean 4 program proven to render every API input correctly, and redesigned every night by an LLM loop that releases without human review whenever a candidate passes. Structure, data fidelity and the no-injection property are proven; contrast, reflow and accessibility are checked per release in a browser and never called proven.
-
Faithful · try it in your browser — Make one TypeScript function faster and see exactly how much of “it still does the same thing” was checked. Each rewrite must compile, stay pure, match the original on generated inputs, survive a bounded Z3 search and benchmark faster; one that is significantly faster is then proved in Lean 4 against a spec agreed in plain words. Early: the translator accepts 39 of 74 corpus functions and 0 of 20 sampled from real libraries.
-
Undefined · try it in your browser — A live program that grows the functions you call but haven’t written. Ask a question of a spreadsheet export; a model drafts the calculation and the checks it must pass, you approve the checks, and a strict TypeScript compiler, unit and property tests, and purity and time-limit gates decide what is accepted. The same engine runs as a CLI and a GitHub Action that certify TypeScript from any source.
-
Dynamic Workflows on Codex — A Claude Code skill: describe a task, and Claude writes a multi-agent workflow script, runs it on Codex agents instead of Claude subagents, and shows the run as a live map.
-
Flightdeck — A Claude Code mod that puts a live agent dashboard in your terminal: context and cost, an advisor timeline, every permission check, and your subagents as cards or swimlanes. It only watches and makes no network requests.
-
nanochat-mlx — Train a small chatbot from scratch on Apple Silicon, from tokenizer training to a chat interface.
-
Qwen Scope Lab — A browser workbench for sparse-autoencoder interpretability on Qwen3.5-2B, running on the Mac through MLX.
ML on Apple Silicon: gemma4-m4-pro · train-gemma4-sudoku-on-your-macbook · ttt-discover-autoresearch-mlx
Research code: bsf-steering · society-of-thought-bench · hypothesis_forge · adaptive_rag_rlm · autoresearch-evo · Proofgrade
macOS menu bar apps (TabPilot, SunShift, SafariMarkdown, GhostLabel, PasteForge, TextDrop and ClipDrop install with brew install scasella/tap/<app>, via homebrew-tap):
TabPilot · SunShift · SafariMarkdown · GhostLabel · PasteForge · TextDrop · ClipDrop · DiskPulse · PortSentry · ProcessBeacon · BrewPilot
- Random-choice adapter — a Qwen3-30B-A3B LoRA for experiments with fixed-list sampling behavior.
- Panel-reasoning adapter — the Qwen3-30B-A3B LoRA from the accuracy and completion-length comparison.
Personal projects. Not affiliated with or endorsed by my employer. Contact: LinkedIn.


