Computing has done this before
From 1946 to 2009, the energy efficiency of computing doubled about every eighteen months. That curve is what moved work from the mainframe to the personal computer. As a recent Stanford study of local AI puts it, the move happened when efficiency let personal devices meet people’s needs within their own power budget, not when PCs beat mainframes in raw performance.1
The same curve is now visible in AI. The Stanford team measured intelligence per watt, task accuracy per unit of power, across more than twenty local models and eight accelerators. From 2023 to 2025 it improved 5.3-fold, and the share of single-turn chat and reasoning queries answered correctly by the best local model of each year rose from 23.2% to 71.3%.1 Models are also getting smaller for the same job: a study in Nature Machine Intelligence found that the capability packed into each parameter of open models doubles about every three and a half months.2
The personal computer did not end the data center. It changed what the data center was for. We expect the same from AI: the cloud keeps the work that needs it, and the computer in front of you takes on more of the rest.
Where AI runs today
In the AI agents most people use today, every step that needs a model runs in the cloud. Take a coding agent. The same frontier model that plans a change also reads every file, interprets every test run and rereads its own history, all at the most expensive point in the stack and inside a context window that Anthropic’s engineers describe as “a finite resource with diminishing marginal returns.”3 It does this not because every step needs a frontier model, but because it has nowhere else to send the work.
Meanwhile, the laptop the agent runs on has a GPU or an NPU that it never uses. And local models are no longer limited to autocomplete or quick answers. The best of them can work as agents in their own right: explore a codebase, make a plan, use tools, run the tests and iterate on what they find. Two capable agents, one in the cloud and one on your desk, and today they don’t work together.
Local AI for everyone
Local AI has come a long way, but it still belongs mostly to enthusiasts: people who can spend thousands on a graphics card or a high-end laptop, and who know their way around model formats, quantization and runtimes. Epoch AI found that a single top gaming GPU, under $2,500, runs open models that match the frontier of six to twelve months earlier.4 That is a high bar for most people.
Most local AI products are built as destinations: another chat window, another runtime, another model to manage. They reward people who enjoy the setup and ask everyone else to change how they work. So the value of local AI stays with the few who invest in hardware and know the technology.
We want the benefits of local AI to reach everyone, on the computer they already own and inside the tools they already use. Nobody should have to pick a runtime, tune a quantization or build a rig to get them.
What we believe
The question is not cloud or local. It is what belongs where. Like computing before it, intelligence will move closer to the people who use it. Frontier models in the cloud will keep doing what only they can do. Capable agents on your own machine will take on the work that lives there, next to your files and your tools. The line between them should follow the nature of the work, not how hard a request looks.
Some of the most underused AI hardware in the world is already on people’s desks. Counterpoint Research expects PCs with a neural processing unit, a chip built for AI, to pass half of global shipments in 2026.5 Most of that capacity plays no part in the AI people use every day. Putting it to work should not be a hobby project.
The best technology disappears into the work. “The most profound technologies are those that disappear,” Mark Weiser wrote in 1991. “They weave themselves into the fabric of everyday life until they are indistinguishable from it.”6 Local AI should work that way, as a quiet part of the tools people already rely on.
Delegation runs on trust. People hand off work when they can see what was done and why. Local intelligence earns its place by showing its work, and by keeping data close to where it lives whenever the work allows.
Research is pointing the same way
NVIDIA researchers argue that small language models are “sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems,” and that agents calling several different models are the natural design when broader abilities are needed.7 Published results back this up. Stanford’s Minions protocol lets a small on-device model read long documents for a cloud model, reducing cloud costs 5.7-fold while keeping 97.9% of the quality on document-reasoning benchmarks.8
Our own early pilot adds a piece about trust: when a local result carried the exact passages it relied on, the cloud agent made the same decisions with far less context.
What we research
We are early, and the deepest questions are still open. That is why Forjal is a lab. Our research follows three themes.
Where local AI fits. Which parts of people’s work belong on their own machine, and which belong in the cloud? We study this inside actual workflows, where the answer matters.
How to get the most from the hardware people own. Consumer laptops combine CPUs, GPUs and NPUs under tight limits of memory and power. We study how to turn that silicon into useful intelligence, machine by machine and model by model.
The system that makes local models capable. A model is only as good as the system around it: how it reads, which tools it can use, how it plans and how it reports back. We are searching for the architecture and the harness that turn every local model into the most capable agent it can be.
Trust, cost and privacy run through all three.
Our first product
Forjal, our first product, brings local intelligence into the agents people already use: Claude Code, Codex, Pi, OpenCode, OpenClaw and Hermes Agent. It runs on macOS and Windows, finds a model that fits your machine and prepares it for you, with no runtime or quantization to choose. Your agent delegates a task to a local agent, which works through it on your machine, inside the folders you allow, and returns the result with the evidence behind it.
There is no new chat to learn and nothing to leave behind: you keep your agent, your projects and the way you already work. And every task the local agent takes on runs on hardware you already own, right where your data lives.
Forjal is in alpha and grows with the research. It is where our questions meet real work.
Where this goes
The personal computer took computing out of the machine room and into everyday life. We believe AI is at the same turning point.
Today, intelligence lives in a handful of data centers and reaches most people through a chat window. We think it will come to live everywhere people work, in the cloud and on the devices they own, cooperating as one system that nobody has to think about.
Forjal is rethinking how people will use AI in everyday life: not as a place they visit, but as something that works alongside them, on the machines they already have. Our first product is where that starts. Our mission is to make local intelligence a natural part of every AI workflow, for everyone.
Your hardware, in the loop.
Sources
-
J. Saad-Falcon, A. Narayan et al., “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI,” Stanford University and Together AI, arXiv:2511.07885, v6, September 2026. The historical trend comes from Koomey et al., IEEE Annals of the History of Computing, as cited by the study, which covers single-turn chat and reasoning queries and query-level routing. ↩ ↩2
-
C. Xiao et al., “Densing law of LLMs,” Nature Machine Intelligence, November 2025, doi:10.1038/s42256-025-01137-0. ↩
-
Anthropic, “Effective context engineering for AI agents,” September 2025, anthropic.com. ↩
-
Epoch AI, “Frontier AI capabilities can be run at home within a year or less,” August 2025, epoch.ai. ↩
-
Counterpoint Research, “AI Advanced PCs to Surpass Half of Global Shipments in 2026 as NPU-powered Laptops Go Mainstream,” September 2025, counterpointresearch.com. ↩
-
M. Weiser, “The Computer for the 21st Century,” Scientific American, September 1991. ↩
-
P. Belcak et al., “Small Language Models are the Future of Agentic AI,” NVIDIA Research, arXiv:2506.02153, 2025. ↩
-
A. Narayan et al., “Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models,” ICML 2025, arXiv:2502.15964. ↩

