Skip to content
@aisa-group

AI Safety and Alignment Group

AI Safety and Alignment Group at the ELLIS Institute Tübingen and Max Planck Institute for Intelligent Systems

ELLIS Institute Tübingen

AI Safety and Alignment Group

ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems

Group page · Updates · Datasets

AI safety · alignment · evaluation

We develop algorithmic approaches to reduce harms from increasingly capable general-purpose AI systems. Our work focuses on the alignment and evaluation of autonomous language-model agents, frontier-model risks and capabilities, and model generalisation and steerability.

Research projects

Project Paper Dataset Website
ResearchArena arXiv Trajectories research-arena.ai
PostTrainBench arXiv Trajectories posttrainbench.com
InferenceBench arXiv Trajectories inferencebench.ai
Instrumental Choices arXiv Agent traces instrumentalchoices.com
Evaluation awareness arXiv EvalAwareBench —
Skill-Inject arXiv — skill-inject.com
Prompt injection in agent skills arXiv — —
QuantSightBench arXiv — quantsightbench.com

Teaching

The AI Safety course at the University of Tübingen is openly available.

Browse all repositories or follow the group on Substack.

Popular repositories Loading

  1. PostTrainBench PostTrainBench Public

    Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

    Python 581 67

  2. tue-ai-safety-course tue-ai-safety-course Public

    AI safety course at the University of Tübingen (Summer Semester 2026)

    HTML 145 12

  3. skill-inject skill-inject Public

    Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

    Python 99 5

  4. InferenceBench InferenceBench Public

    Benchmarking Open-Ended Inference Optimization by AI Agents

    Python 47 8

  5. promptinject-agent-skills promptinject-agent-skills Public

    Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections

    Python 23 4

  6. decomposing-eval-awareness decomposing-eval-awareness Public

    Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.

    Python 20 4

Repositories

Showing 10 of 19 repositories
  • posttrainbench-website Public

    Website for PostTrainBench

    aisa-group/posttrainbench-website's past year of commit activity
    JavaScript 2 1 0 0 Updated Oct 2, 2026
  • PostTrainBench Public

    Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

    aisa-group/PostTrainBench's past year of commit activity
    Python 581 MIT 67 12 10 Updated Oct 2, 2026
  • aisa-group/training_against_probes's past year of commit activity
    Python 4 1 0 0 Updated Oct 2, 2026
  • aisa-group/instrumental-evasion's past year of commit activity
    Python 4 0 0 0 Updated Sep 30, 2026
  • decomposing-eval-awareness Public

    Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.

    aisa-group/decomposing-eval-awareness's past year of commit activity
    Python 20 4 0 0 Updated Sep 29, 2026
  • perfect-crime Public
    aisa-group/perfect-crime's past year of commit activity
    Python 6 3 0 0 Updated Sep 25, 2026
  • ResearchArena Public
    aisa-group/ResearchArena's past year of commit activity
    Python 7 MIT 2 0 0 Updated Sep 25, 2026
  • aisa-group/inferencebench-site's past year of commit activity
    JavaScript 0 0 0 0 Updated Sep 25, 2026
  • skill-inject Public

    Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

    aisa-group/skill-inject's past year of commit activity
    Python 99 MIT 5 1 0 Updated Aug 29, 2026
  • tue-ai-safety-course Public

    AI safety course at the University of Tübingen (Summer Semester 2026)

    aisa-group/tue-ai-safety-course's past year of commit activity
    HTML 145 12 0 0 Updated Aug 20, 2026