Skip to content
@aisilab

AI Safety & Interpretability Lab

Popular repositories Loading

  1. arbiter arbiter Public

    Run HuggingFace models through freeform questions and judge responses with an LLM.

    Python 5 1

  2. psychological-safety psychological-safety Public

    Python 2

  3. rapidrouge rapidrouge Public

    Drop-in rouge-score replacement with a bit-parallel ROUGE-L

    Python 2

  4. diffing-toolkit diffing-toolkit Public

    Forked from science-of-finetuning/diffing-toolkit

    A toolkit that provides a range of model diffing techniques including a UI to visualize them interactively.

    Python 1 1

  5. aisilab.github.io aisilab.github.io Public

    Website of the AI Safety & Interpretability Lab at SDU

    HTML 1 1

  6. model_organism_trainer model_organism_trainer Public

    One-file Unsloth LoRA finetuner for taboo model organisms

    Python 1

Repositories

Showing 10 of 13 repositories

Top languages

Loading…

Most used topics

Loading…