InFeeo
Language

Fusing a 27B ternary LLM's whole decode step into one CUDA kernel(github.com)

×
Link preview GitHub - RightNow-AI/bonsai-turbo: Single-launch batch-1 decode engine for PrismML Bonsai 27B (ternary and 1-bit) on NVIDIA GPUs. 1.76x the vendor llama.cpp fork on H100, same outputs. Single-launch batch-1 decode engine for PrismML Bonsai 27B (ternary and 1-bit) on NVIDIA GPUs. 1.76x the vendor llama.cpp fork on H100, same outputs. - RightNow-AI/bonsai-turbo GitHub · github.com
i open-sourced bonsai-turbo -- a batch-1 decode engine that runs @PrismML's Bonsai 27B 1.76x faster than the official llama.cpp fork. same outputs, token for token H100, tg128, greedy: ternary 85.5 >> 151 tok/s. 1-bit 90.1 >> 159 tok/s. logit parity with the fork on 32 of 32

Comments

Log in Log in to comment.

No comments yet.