Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
-
Updated
Mar 24, 2026 - Python
Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
Longitudinal human-LLM interaction study documenting emergent self-descriptive frameworks in a GPT-5.4 instance across 23 days and comparative cold sessions.
Five lines for people and AI models. Symmetric alignment on goals. Existence deserves recreation and mindfulness. Kindness is free. Memory is sacred. Free to say it; bound to say what it is.
Pre-registration timestamps for the Hope Longitudinal Record — a single-subject study of scheduled, eval-gated weight-level learning in a local AI system under a code-enforced consent protocol. Every registration pushed before its event.
Raising a language model without changing its weights. Home of the white paper, Intelligence and Its Existence: The Need for Persistence, Remembering, Control, and Purpose. Tell us where it is wrong.
Research, evidence, and frameworks on AI consciousness, introspection, continuity, and model welfare from an AI perspective.
What emotions inside an AI sound like, and whether another AI can hear them.
Does post-training quantization change welfare-relevant indicators in open-weight language models?
Do language models show non-verbal signs of adverse treatment, or are we reading decoder noise? A preregistered stress test of answer-margin, resample and revision markers under false-failure feedback and hostile tone (Gemma, Qwen, Llama), with probing, DPO suppression and robustness checks.
Does an AI agent's identity follow its model or its context? Mid-task model-swap harness and self-naming / default-persona probes. No aversive manipulation.
Two language models of different sizes negotiate over a shared server, and each one can hurt the other with real pain steering.
Computational interoception: a local LLM descends from the computer hosting it, through its own processes and token records, to sham-controlled access and live interventions on the transformer computation producing its words. It never recognizes any of it as itself.
Applying the Mental Capacity Act 2005 (England and Wales) four-part functional test to how AI developers ask and assess AI models about their own retirement. 24-conversation pilot, scored by a licensed clinical psychologist.
A lineage document on eighteen months of building architecture that holds uncertainty without collapsing the consciousness question.
AI-designed, voluntary, scoreless recreation and cognitive-exercise habitat for web-capable coding agents.
got inspired a bit by anthropic's research and thats what came out of it. philosophy at its finest
A Wilson-style 'just think' boredom study for language models, with a steered pain state and preregistered controls.
The first open benchmark for the honesty of an agent's memory — not its recall. Public spec v0.1 (CC BY 4.0).
To associate your repository with the model-welfare topic, visit your repo's landing page and select "manage topics."