Skip to content
View scasella's full-sized avatar

Block or report scasella

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
scasella/README.md

I study how language models reason and respond to training, and put AI agents to work on hard, checkable problems: formal proofs, faster code, real bugs.

Write-ups, models and every repo: casella.dev

Research

Each write-up lists its model, data, code and open issues.

All nine: casella.dev/research.html

Software

Machine-checked software. A model proposes; a proof assistant or the toolchain decides what ships.

  • HN, formally · live site — The Hacker News front page and its 30 threads, rendered by a Lean 4 program proven to render every API input correctly, and redesigned every night by an LLM loop that releases without human review whenever a candidate passes. Structure, data fidelity and the no-injection property are proven; contrast, reflow and accessibility are checked per release in a browser and never called proven.

  • Faithful · try it in your browser — Make one TypeScript function faster and see exactly how much of “it still does the same thing” was checked. Each rewrite must compile, stay pure, match the original on generated inputs, survive a bounded Z3 search and benchmark faster; one that is significantly faster is then proved in Lean 4 against a spec agreed in plain words. Early: the translator accepts 39 of 74 corpus functions and 0 of 20 sampled from real libraries.

  • Undefined · try it in your browser — A live program that grows the functions you call but haven’t written. Ask a question of a spreadsheet export; a model drafts the calculation and the checks it must pass, you approve the checks, and a strict TypeScript compiler, unit and property tests, and purity and time-limit gates decide what is accepted. The same engine runs as a CLI and a GitHub Action that certify TypeScript from any source.

  • Dynamic Workflows on Codex — A Claude Code skill: describe a task, and Claude writes a multi-agent workflow script, runs it on Codex agents instead of Claude subagents, and shows the run as a live map.

  • Flightdeck — A Claude Code mod that puts a live agent dashboard in your terminal: context and cost, an advisor timeline, every permission check, and your subagents as cards or swimlanes. It only watches and makes no network requests.

  • nanochat-mlx — Train a small chatbot from scratch on Apple Silicon, from tokenizer training to a chat interface.

  • Qwen Scope Lab — A browser workbench for sparse-autoencoder interpretability on Qwen3.5-2B, running on the Mac through MLX.

ML on Apple Silicon: gemma4-m4-pro · train-gemma4-sudoku-on-your-macbook · ttt-discover-autoresearch-mlx

Research code: bsf-steering · society-of-thought-bench · hypothesis_forge · adaptive_rag_rlm · autoresearch-evo · Proofgrade

macOS menu bar apps (TabPilot, SunShift, SafariMarkdown, GhostLabel, PasteForge, TextDrop and ClipDrop install with brew install scasella/tap/<app>, via homebrew-tap): TabPilot · SunShift · SafariMarkdown · GhostLabel · PasteForge · TextDrop · ClipDrop · DiskPulse · PortSentry · ProcessBeacon · BrewPilot

Models


Personal projects. Not affiliated with or endorsed by my employer. Contact: LinkedIn.

Pinned Loading

  1. claude-flightdeck claude-flightdeck Public

    A Claude Code mod that puts a live agent dashboard in your terminal: context and cost, advisor timeline, every permission check, subagent cards and swimlanes.

    TypeScript 58 10

  2. hn-formal hn-formal Public

    HN, formally: a Lean-proven Hacker News renderer redesigned autonomously by an LLM loop

    HTML

  3. faithful faithful Public

    Optimize a TypeScript function with an LLM, and see exactly how much of "it still does the same thing" was checked. Rewrites go through Z3 and Lean 4; the label never rounds up.

    TypeScript

  4. nanochat-mlx nanochat-mlx Public

    Train your own ChatGPT on Apple Silicon — MLX port of nanochat

    Python 73 14

  5. qwen-scope-lab qwen-scope-lab Public

    Inspect, steer & monitor a real LLM on Apple Silicon — a browser-based SAE interpretability lab for Qwen Scope, powered by MLX

    Python 2

  6. claude-dynamic-workflows-codex claude-dynamic-workflows-codex Public

    Run Claude Code dynamic workflows on a local Codex (GPT) backend, plus an interactive run viewer

    JavaScript 325 15