Speculative Decoding Performance Analyzer
7.5
A tool to analyze and debug LLM inference performance issues by reproducing speculative decoding setups and measuring their efficiency, identifying bottlenecks related to hardware or model configurations. Facilitates fine-tuning speculative decoding parameters.
120h
mvp estimate
7.5
viability grade
0
views
technology stack
Python
Medium
inspired by
Speculative Decoding's Speedup a Hardware Problem or a Model Problem?