Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 34 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# Changelog




## linghe 0.3.0

- use faster and more accurate implementation for `embedding` backward
- add multiple embedding_lookup implementations
- use 2048 block number for `batch_count_zero` kernel


## linghe 0.2.9

- support stride in grad tensor of `embedding` kernel
- return grad for dummy tensor in `embedding` kernel
- support bf16 in batch mul/clip/norm kernels


## linghe 0.2.8

- fix racing condition bug in softmax_cross_entropy kernel
- use tl.rsqrt instead of 1/tl.sqrt in all kernels
- add the parameter `tp_group` to `softmax_cross_entropy`


## linghe 0.2.7

- add the parameter `ignore_index` to `softmax_cross_entropy`
- support parallel `softmax_cross_entropy`
- add dtype and numel assertion in multiple batch kernels

- Known issues:
- performance of `softmax_cross_entropy` degrades when vocab size is not multiple of 16
14 changes: 12 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@

## Introduction
---
Our repo, linghe, is designed for LLM training, especially for MoE training with FP8 quantizaiton. It provides 2 main categories of kernels:
Our repo, linghe, is designed for LLM training, especially for MoE training with FP8 quantizaiton. It provides 3 main categories of kernels:

- **Fused quantization kernels**: fuse quantization with previous layer, e.g., RMS norm and Silu.
- **Memory-efficiency kernels**: fuse multiple IO-itensive operations, e.g., ROPE with qk-norm.
Expand Down Expand Up @@ -66,4 +66,14 @@ Examples can be found in tests.
## Api Reference
---

Please refer to [API](https://inclusionai.github.io/linghe/)
Please refer to [API](https://inclusionai.github.io/linghe/)

## Citations

[TBD]
```
@misc{zhao2025linghe,
title={Linghe: Enabling Efficient Trillion-Scale LLM Training via Optimized Kernels},
author={Yao Zhao and Chen Liang and Jingyu Hu and Zixuan Cheng and Longfei Li}
}
```
Loading