AI Model Evaluation | RLHF | Prompt Engineering. Focused on adversarial testing, constraint satisfaction, and creating rigorous evaluation rubrics.
Pinned Loading
-
llm-logical-integrity-benchmark
llm-logical-integrity-benchmark PublicAdversarial testing of LLMs on constraint satisfaction deadlocks
-
rag-evaluation-with-ragas
rag-evaluation-with-ragas PublicEvaluation-driven RAG system using RAGAS, retrieval metrics, behavioral stress testing, holdout validation, diagnostics and reliability analysis.
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.