Exploring adversarial robustness, uncertainty, and model awareness through experiments, notebooks, and technical essays.
This repository is not simply a collection of notebooks.
It is a record of an evolving investigation.
Each notebook accompanies an essay.
Each essay begins with a question.
Each question changes the next experiment.
Some conclusions may change.
Some assumptions may fail.
That is part of the process.
Do not ask the experiment to confirm the question.
Ask it to change the question.
The series began with a simple question:
How are we vulnerable?
From there, the investigation moved through adversarial perturbations, defenses, uncertainty, out-of-distribution detection, decision-boundary geometry, and architectural awareness.
The broader question gradually became:
Can a model become more aware of when its own predictions may be unreliable?
The goal of this repository is not only to reproduce results, but to document the process of inquiry — including experiments, failures, changing assumptions, and unexpected directions.
This series is designed as a cumulative investigation rather than a collection of isolated experiments.
Its value lies in connecting several reliability questions that are often studied separately:
- adversarial vulnerability;
- defensive robustness;
- uncertainty estimation;
- out-of-distribution detection;
- selective prediction;
- decision-boundary geometry;
- architectural uncertainty signals.
Each essay and notebook tests a question, records what worked and what failed, and uses that evidence to motivate the next stage of the series.
The repository therefore serves as:
- a reproducible record of the experimental path;
- a technical companion to the published essays;
- a comparison of different signals for model uncertainty and adversarial risk;
- a study of how reliability mechanisms behave under both natural and adversarial conditions;
- a foundation for future work on model awareness and trustworthy AI.
The broader contribution of the series is not a claim that one mechanism solves adversarial robustness, but a structured attempt to understand where different forms of awareness succeed, fail, and complement one another.
If you are new to the series, I recommend reading the essays in order:
-
Essay #1 — The Fog of Certainty: On Deep Learning's Secret Vice
Published on Medium -
Essay #2 — The Attack Landscape: Why Silence Is the Vulnerability
Published in Towards AI -
Essay #3 — Walls, Shields, and Illusions: Defenses and Their Limits
Published in Towards AI -
Bridge Essay — The Bridge: From Resistance to Awareness — Why Defenses Alone Will Never Be Enough
Published on Medium -
Essay #4 — Learning to Say “I Don’t Know”: The First Floor of Awareness
Published in Towards AI -
Essay #5 — The Geometry of Fragility: Feeling the Decision Boundary
Published on Medium -
Essay #6 — Architectural Awareness: Building the Sensor into the Boat
Published on Medium -
Essay #7 — Knowing the Weights: Four Bayesian Lenses on Uncertainty
Published on Medium -
Essay #8 — Does the Model Know Where It Looks
In progress -
Essay #9 _
Planned
The attacker adapts. The wall crumbles.
The researcher adapts. The question changes.
This repository accompanies an ongoing study of adversarial attacks, defenses, uncertainty, and model behavior.
The series is intentionally cumulative: each experiment exposes a limitation that motivates the next question.
The starting point of the series.
The essay begins with a shift in framing:
Not “Are we vulnerable?” — but “How are we vulnerable?”
It explores adversarial vulnerability, assumptions about robustness, and the foundations of the questions developed throughout the series.
📖 Article
Read on Medium
Notebook
- No companion notebook for this essay.
A study of adversarial behavior beyond visible failures, asking whether omission, selective behavior, and silence can themselves become vulnerabilities.
📖 Article
Read in Towards AI
Notebook
An exploration of adversarial training, defensive distillation, gradient masking, and a deeper question:
Can any defense remain unbreakable given enough computational power and knowledge of the model?
If every wall eventually breaks, then the entire framework of defense begins to look incomplete.
📖 Article
Read in Towards AI
Notebook
Why Defenses Alone Will Never Be Enough
The bridge asks a different question.
Instead of:
How can we build a stronger wall?
it asks:
What if resistance alone is the wrong objective?
This essay connects adversarial robustness with a broader perspective: model awareness.
It is not the conclusion of the first part of the journey.
It is the path leading to the next one.
📖 Article
Read on Medium
Notebook
- No companion notebook for this essay.
We built the voice.
A decision gate — a simple, auditable rule that says:
“If uncertainty is high, do not predict. Ask for help.”
This is not a new wall.
It is a voice, and it is the first floor of a structure this series will keep building on.
The experiment showed a strong ability to defer on natural out-of-distribution data, while exposing a major limitation under adversarial perturbation.
📖 Article
Read in Towards AI
Notebook
We gave the model a voice — a way to say:
“I am not sure.”
and
“This input looks strange.”
It worked on natural data.
But it failed on adversarial data.
That was the blind spot we needed to close.
The model needed to become aware not only of the data distribution, but also of its own decision geometry.
This essay explores boundary proximity as a signal of fragility.
📖 Article
Read on Medium
Notebook
The architecture was still designed primarily to produce predictions.
Awareness had been added afterward through external measurements and decision gates.
Now the question changed:
Can we design the boat itself to be more aware?
Can uncertainty become part of the architecture rather than an external attachment?
This essay explores two directions:
- Multiple Prediction Heads — a shared backbone with several prediction heads whose disagreement becomes an internal uncertainty signal.
- Evidential Outputs — a model whose output contains not only a prediction, but also an estimate of the evidence supporting that prediction.
The results show that internal architectural signals can outperform plain confidence under adversarial evaluation, while also revealing the limits of combining signals naively.
📖 Article
Read on Medium
Notebook
The Boat and Its Weights!
The weights are where the model stores its learned representation of the world. They are the memory of its training experience.
They determine how it transforms an input into a prediction.
Can the model know whether its own weights are reliable?
This is the question explored in Essay #7.
📖 Article
Read on Medium
Notebook
The Gaze of the Boat.
what does it mean for a signal to be decoupled from failure?
Before measuring attention, we need to be precise about what the experiment does not claim.
An attention map can look like a picture of the model’s focus, but a readable mechanism is not automatically
a faithful explanation of the model’s decision.
📖 Article
Read on Medium
Notebook
Status: in progress
Planned future direction.
This part of serries will explore What would it mean for a system to be honest about its own vulnerability?
Status: Planned
- Natural OOD deferral: 80.09%
- Adversarial deferral: 8.80%
This experiment showed that uncertainty and OOD detection could identify unfamiliar natural inputs while still failing to recognize many adversarially perturbed inputs.
| Signal | Adversarial Risk |
|---|---|
| Confidence | 78.70% |
| Multi-head disagreement | 19.86% |
| Evidential evidence | 21.31% |
| Boundary distance | 77.06% |
| Fusion | Adversarial Risk |
|---|---|
| Without boundary | 60.61% |
| With boundary | 60.21% |
These results suggest that multi-head disagreement and evidential signals provided substantially stronger adversarial-risk separation than confidence or boundary distance alone under the fair comparison setup.
The series currently explores:
- adversarial attacks
- adversarial defenses
- uncertainty estimation
- out-of-distribution detection
- selective prediction
- decision-boundary geometry
- architectural uncertainty
- evidential deep learning
- model awareness
- trustworthy and transparent ML experimentation
- Python
- PyTorch
- NumPy
- Matplotlib
- scikit-learn
- Jupyter Notebook
- MNIST
- Fashion-MNIST
- CIFAR-10
- ViT
- FGSM
- PGD
- Monte Carlo Dropout
- Deep Ensembles
- SWAG
- Variational Inference
- Attention entropy
/
├── #2/
├── #3/
├── #4/
├── #5/
├── #6/
├── #7/
├── #8/
├── docs/
├── data/
├── .gitattributes
└── README.md
- Not knowing is not the opposite of research; Not knowing is where research begins.*
The experiments in this repository are research-scale studies rather than production evaluations.
Important limitations include:
- most experiments use benchmark datasets such as MNIST, Fashion-MNIST, and CIFAR-10;
- conclusions may not transfer directly to larger models or real-world domains;
- adversarial evaluation depends on the selected attacks, threat models, and hyperparameters;
- some awareness signals remain computationally expensive or difficult to compare fairly;
- results should be interpreted as evidence within the experimental setup, not as universal guarantees;
- later essays may revise or challenge earlier interpretations as the investigation develops.
These limitations are intentional parts of the research process and are documented rather than hidden.
This repository evolves alongside the essays.
Some conclusions may change.
Some assumptions may fail.
Some questions may become more interesting than their answers.
That is part of the process.
-
“The road has been changed in every movement.”
-
“Every honest question builds the next bridge.”
Applied AI & Machine Learning Engineer exploring adversarial robustness, uncertainty, and model awareness through experiments and technical writing.
