Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 23 additions & 1 deletion .github/workflows/pages.yml
Original file line number Diff line number Diff line change
Expand Up @@ -38,9 +38,31 @@ jobs:
run: |
mkdir -p _site
cp index.html 404.html robots.txt sitemap.xml .nojekyll _site/
cp -r assets legal ru _site/
cp -r assets checks legal ru _site/
find _site -type f | sort

# Deliberate naming above means a forgotten directory is silently absent
# from the published site rather than from the repository: /checks/ was
# committed, passed every structural check, and 404ed for twenty minutes
# while /ru/checks/ worked, because `ru` is copied whole and `checks` was
# not on the list. check_site.py knows which pages exist; this asks it.
- name: Every page the checker knows about is actually published
run: |
python - <<'PY'
import pathlib, re, sys
source = pathlib.Path("scripts/check_site.py").read_text(encoding="utf-8")
match = re.search(r"^PAGES = \[(.*?)\]", source, re.MULTILINE | re.DOTALL)
if not match:
sys.exit("check_site.py no longer declares PAGES in the expected form")
pages = re.findall(r'"([^"]+)"', match.group(1))
missing = [p for p in pages if not (pathlib.Path("_site") / p).is_file()]
for page in missing:
print(f" not published: {page}")
if missing:
sys.exit(f"{len(missing)} page(s) in the repository never reach the site")
print(f" all {len(pages)} pages published")
PY

- uses: actions/configure-pages@v6
- uses: actions/upload-pages-artifact@v5
with:
Expand Down
16 changes: 11 additions & 5 deletions checks/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -150,6 +150,11 @@ <h2>Numbers pinned to code</h2>
<td>praxis README, both languages</td>
<td><code>tests/test_eval.py</code></td>
</tr>
<tr>
<td>praxis's offline retrieval figures are what a run actually produces</td>
<td>the quality table in both praxis READMEs</td>
<td><code>tests/test_eval.py</code>, under <code>PRAXIS_OFFLINE=1</code></td>
</tr>
<tr>
<td>praxis depends on this organisation's own packages, not on strangers' names</td>
<td><code>pyproject.toml</code> extras</td>
Expand Down Expand Up @@ -242,12 +247,13 @@ <h2>What nothing checks yet</h2>

<dl class="facts" style="max-width:44rem;margin-top:20px">
<div>
<dt>praxis retrieval metrics</dt>
<dt>praxis on real models</dt>
<dd>
recall@5 0.92 and MRR 0.94 on the full corpus are stated in prose. They depend on
which models are installed, so holding them needs a scheduled re-run like
decisionrl's, not a unit test. The golden set's size and a recall floor of 0.7 are
pinned; the headline figures are not.
recall@5 0.92 and MRR 0.94 on the full corpus are stated in prose and measured on a
GPU that CI does not have. The offline column of the same table is now held by a test
— the offline path has no models, no network and no seed, so a run of it is
reproducible — but the GPU column needs a scheduled re-run on real hardware, the way
decisionrl re-verifies its applied claims nightly.
</dd>
</div>
<div>
Expand Down
16 changes: 11 additions & 5 deletions ru/checks/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -151,6 +151,11 @@ <h2>Числа, прибитые к коду</h2>
<td>README praxis, обе версии</td>
<td><code>tests/test_eval.py</code></td>
</tr>
<tr>
<td>Офлайн-метрики поиска praxis — это то, что печатает прогон</td>
<td>таблица качества в обоих README praxis</td>
<td><code>tests/test_eval.py</code>, под <code>PRAXIS_OFFLINE=1</code></td>
</tr>
<tr>
<td>praxis зависит от пакетов этой организации, а не от чужих имён</td>
<td>extras в <code>pyproject.toml</code></td>
Expand Down Expand Up @@ -242,12 +247,13 @@ <h2>Что пока не проверяется</h2>

<dl class="facts" style="max-width:44rem;margin-top:20px">
<div>
<dt>Метрики поиска praxis</dt>
<dt>praxis на реальных моделях</dt>
<dd>
recall@5 0.92 и MRR 0.94 по полному корпусу заявлены прозой. Они зависят от того,
какие модели установлены, поэтому держать их нужно перезапуском по расписанию, как в
decisionrl, а не юнит-тестом. Размер золотого набора и нижняя граница recall 0.7
прибиты; сами заголовочные числа — нет.
recall@5 0.92 и MRR 0.94 по полному корпусу заявлены прозой и замерены на GPU,
которого в CI нет. Офлайн-колонка той же таблицы теперь держится тестом — у офлайн-пути
нет ни моделей, ни сети, ни случайности, поэтому его прогон воспроизводим, — а колонка
с реальными моделями требует перезапуска по расписанию на настоящем железе, как
decisionrl перепроверяет свои прикладные заявления еженощно.
</dd>
</div>
<div>
Expand Down
Loading