What is it about?
Running a large AI language model on a personal device is difficult because its data can exceed the device’s working memory. Keeping that data in flash storage helps it fit, but repeatedly moving it to a processor can make responses very slow. We propose Cambricon-FlexLLM, a hardware design that lets flash memory do some of the calculations alongside an AI processor. It also groups model data that is often used together and adapts the workload to avoid transferring data that is not needed for the current calculation. In simulations, the design generates 3.44 tokens—small pieces of text—per second with a 70-billion-parameter model. Further optimizations that exploit inactive parts of the model provide an average 1.7-fold speedup across the tested models and hardware configurations.
Featured Image
Photo by Solen Feyissa on Unsplash
Why is it important?
Local AI could help people use capable assistants when internet access is unreliable or when sensitive information should stay on their devices. This work tackles a key obstacle: moving the large amount of data these models need. Its distinctive contribution is to coordinate an AI processor with flash memory that can both store data and perform calculations, while adapting data placement and work sharing to the parts of the model being used. The results point toward local use of very large language models on future personal devices. Practical deployment still requires further work on prompt-processing delays, power, heat, and storage reliability.
Perspectives
What I find most compelling about this work is the possibility of giving people more choice over where their AI runs. For a private question or a task without reliable internet access, having a local option could make a real difference. I hope this paper encourages researchers to rethink how storage and processors work together, and to consider data movement as a central part of making large AI models useful on personal devices.
tianyun ma
Institute of Industrial Artificial Intelligence, Chinese Academy of Sciences
Read the Original
This page is a summary of: Cambricon-FlexLLM: A Flexible Chiplet-Based Hybrid Architecture for On-Device 70B LLM Inference, ACM Transactions on Architecture and Code Optimization, August 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3844618.
You can read the full text:
Contributors
The following have contributed to this page







