about
Rotary GPU: Exploring Local Execution for Large MoE Models Under Limited VRAM (arxiv.org)
41 points by dryarzeg 126 days ago | hide | past | pdf | 4 comments on HN

In plain words: Rotary GPU keeps only the parts of a large expert-based language model needed in the graphics card's memory, so it runs on laptops. On an 8 GB laptop card it wrote 2,048 tokens at 21 per second using 6.3 GB, not a big server.

Abstract · Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory

Large language models have achieved remarkable capabilities through scaling, and this paper does not challenge that. It instead investigates a different question: once large models already exist, can they become more accessible to environments with substantially smaller hardware resources? The motivation came from deployment concerns rather than architecture research. Many organizations operate under hardware, budget, security, or closed-network constraints that limit access to large accelerator clusters, and as models continue to improve, deployment accessibility may matter as much as capability itself. This paper presents Rotary GPU, an exploratory execution approach derived from a previously disclosed rotary-based accelerator residency concept. A public validation was conducted using a Qwen3.6-35B-A3B-class Mixture-of-Experts model executed locally on a consumer laptop with an RTX 4060 Laptop GPU containing 8 GB of VRAM. Under the primary configuration, the system generated 2048 output tokens while maintaining approximately 6.3 GB of VRAM usage and an observed decode throughput of 21.06 tokens per second. The goal is not to replace data-center infrastructure but to explore whether some capabilities of large models can be brought closer to environments where such infrastructure is unavailable. The results should be read as exploratory rather than definitive, but they suggest deployment accessibility deserves continued investigation as these models evolve.

Myeong Jun Jo
arXiv:2605.29135 · cs.PF, cs.AR, cs.DC · submitted May 27, 2026
abstract · pdf · html · 10 pages, 3 figures. Also archived at Zenodo (DOI: 10.5281/zenodo.20406471). Related to Korean Patent Publication KR 10-2026-0070380

add comment on HN

Why is this a paper? It's just using the n-cpu-moe option on llama.cpp? What am I missing here?
It's amazingly vacuous isn't it? I think the most interesting read was the fact that they were surprised llama.cpp crashed when they used a bad set of commandline arguments.

Although in the section immediately above the observation they claimed that they ran 10 whole completions with 100% success rate. So who knows.

I have to admit I slightly miss the flood of AI-psychosis research papers that seemed to be popping up a couple of months ago. Good to know there's still one or two new ones floating around.

Apparently the author has a patent about it, too.
Um, doesn't the 4060 laptop card have the ability to share system memory?

Wait... My mistake. Google AI says the 4060 mobile can access system memory but tech sheets say no.