-
Notifications
You must be signed in to change notification settings - Fork 211
Pull requests: lightseekorg/tokenspeed
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
refactor(kernel): organize amd gluon kernels based on generations
#880
opened Aug 1, 2026 by
borontion
Contributor
Loading…
[WIP](perf): Optimize Kimi K3 MoE input projections and low-token decode
#878
opened Jul 31, 2026 by
panditsa
Contributor
Loading…
feat(amd): add initial GFX1250 Gluon MLA kernels
#877
opened Jul 31, 2026 by
Yu-Zhewen
Contributor
Loading…
fix(kimi-k3): serve on GB200 -- LCM packing, layer derivation, kda la…
#876
opened Jul 31, 2026 by
nperrin-fr
Collaborator
•
Draft
Remove Radix cache path and consolidate runtime cache handling
#864
opened Jul 31, 2026 by
wangbo981016
Contributor
Loading…
Support detailed token logprobs for RL and distillation
#861
opened Jul 31, 2026 by
HJSang
Collaborator
Loading…
perf(comm): MNNVL fused all-reduce (one-shot + two-shot) for cross-node TP
#860
opened Jul 31, 2026 by
nperrin-fr
Collaborator
Loading…
[WIP] perf(kda): CuTe DSL fused decode for KDA linear attention
#857
opened Jul 30, 2026 by
nperrin-fr
Collaborator
•
Draft
perf(kimi-k3): fuse AMD native MoE up projection epilogue
#841
opened Jul 29, 2026 by
qedawkins
Contributor
Loading…
perf: fuse MiniMax sparse cache insertion
#839
opened Jul 29, 2026 by
FlamingoPg
Contributor
•
Draft
fix(runtime): keep multimodal tensors inline across nodes
#826
opened Jul 28, 2026 by
tuanzhangCS
Contributor
•
Draft
deps: test TensorRT-LLM rc22 kernel wheel
#818
opened Jul 27, 2026 by
Xiangyi1996
Collaborator
•
Draft
feat(kernel): support small-batch Gluon MLA decode
#793
opened Jul 24, 2026 by
Max191
Contributor
Loading…
Previous Next
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.