O
Jul 4, 2026One Model, Three Chips, Two Files: How LiteRT Delegates Really Work
Three chips can run the same LLM, but it takes two files - one shared by CPU and GPU, one just for the NPU. Understanding why took me into graphs, subgraphs, and the difference between just-in-time and ahead-of-time compilation - the layer LiteRT calls delegates.
Jul 4, 202610 min read0 reactions0 comments


