Skip to content
View yinli-systems's full-sized avatar

Block or report yinli-systems

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yinli-systems/README.md

Yin Li

ML systems engineering · GPU inference · efficient training

I build and profile model runtimes, training pipelines, and evaluation infrastructure. My work connects a concrete bottleneck or correctness failure to a small implementation, reproducible evidence, and upstream review.

Portfolio · Upstream contributions · Project index · Contribution workspaces

Selected upstream contributions

Project Contribution Evidence
vLLM Generate output logprobs now carry integer token IDs directly instead of "token_id:N" placeholder strings, with aligned Python/Rust frontend behavior, rank preservation, derender compatibility, tests, and documentation. Merged #58181
PyTorch TorchTitan Vocabulary-sharded RL policy statistics: avoid full-vocabulary gathers while preserving logprob/gradient semantics and the batch-invariant fallback. Merged #4685
Apache DataFusion Authored: make sliced list set operations process the visible range rather than hidden backing data; one documented component case improved from 184.46 µs to 0.823 µs (not a whole-query speedup). Reviewed: merged Parquet RLE/dictionary-read PR #24227; benchmark analysis identified up to ~9× regressions on high-cardinality GROUP BY workloads, and the findings were cited by the PR author in follow-up performance discussion. Merged #25300 · Reviewed, merged #24227
Apache Arrow-rs Harden temporal casts against overflow: respect safe/unsafe cast semantics for narrowing Time64, use checked precision increases, and add regression coverage for remaining temporal conversion paths. Merged #10907
Triton Incomplete-cache recovery, removal of unnecessary single-config autotuning, and tensor-indexing correctness. Merged #10411, #10413, #10429
Inspect AI Bounded continuation of oversized sandbox JSON-RPC responses, with transport validation and cleanup. Merged #4371
GPU attention runtimes CuTe SM120/SM121 compile-time handling in FlashAttention; per-kernel large-head shared-memory initialization in ONNX Runtime. Merged FlashAttention #2671, ONNX Runtime #29140

Also: the SWE-bench skipped-test grading fix was adopted through upstream commit a5ecda6. PR #598 is closed, not marked merged. The full record keeps authored PRs, adopted changes, reviews, and reports separate.

Flagship projects

Project What to inspect
Single-GPU Inference Lab Request-level prefill cost geometry, full-engine profiling, and measured serving changes. Raw artifacts, cross-hardware controls, and negative results accompany the claims. Start with the report.
L20 Pretraining Lab A released 1.100B language model trained from scratch on 20.0B prediction tokens on one L20, with frozen evaluation records. Multimodal work is a separate research track, not a promoted all-purpose release. Language model card.
L20-CodeForge Executable-code evaluation and verifier-guided inference, with full-suite artifacts, public/hidden-test separation, and a separate model-weight retention audit. Reproduction guide.

Selected work under review

vLLM sampled-token logprob transport — #57442 recovers measured L20 serving throughput from 0.67× to 0.95–0.96× of the no-logprobs baseline on the documented Qwen2.5-0.5B workloads. Open, not merged. More open work and its validation boundaries.

Python · C++/CUDA · Rust · Triton · PyTorch · vLLM

@yinli-systems's activity is private