about
An open-source energy and latency profiler for LLMs: ELANA (arxiv.org)
1 point by nkko 294 days ago | hide | past | pdf | discuss on HN

In plain words: ELANA is a free, lightweight tool that measures how big a language model is, how much memory its saved conversation history needs, how fast it produces its first and later words, and how much power it uses. It runs on both multi-GPU servers and small edge chips, and works with every public model on Hugging Face through a simple command line.

Abstract · ELANA: A Simple Energy and Latency Analyzer for LLMs

The latency and power consumption of large language models (LLMs) are major constraints when serving them across a wide spectrum of hardware platforms, from mobile edge devices to cloud GPU clusters. Benchmarking is crucial for optimizing efficiency in both model deployment and next-generation model development. To address this need, we open-source a simple profiling tool, \textbf{ELANA}, for evaluating LLMs. ELANA is designed as a lightweight, academic-friendly profiler for analyzing model size, key-value (KV) cache size, prefilling latency (Time-to-first-token, TTFT), generation latency (Time-per-output-token, TPOT), and end-to-end latency (Time-to-last-token, TTLT) of LLMs on both multi-GPU and edge GPU platforms. It supports all publicly available models on Hugging Face and offers a simple command-line interface, along with optional energy consumption logging. Moreover, ELANA is fully compatible with popular Hugging Face APIs and can be easily customized or adapted to compressed or low bit-width models, making it ideal for research on efficient LLMs or for small-scale proof-of-concept studies. We release the ELANA profiling tool at: https://github.com/enyac-group/Elana.

Hung-Yueh Chiang, Bokun Wang, Diana Marculescu
arXiv:2512.09946 · cs.DC, cs.AI · submitted Dec 7, 2025
abstract · pdf · html

add comment on HN