Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.
Turn any computer or edge device into a command center for your computer vision projects.
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.