The hardware is already here.
The hardware for local AI is already on our desks. CPUs, GPUs, and NPUs in ordinary computers can run models capable of real, bounded work. The capacity is uneven, as it should be: each machine carries a different load, and each model has different strengths. Making that capacity useful starts with honesty about the differences: what is merely available, what actually runs on a machine, and what has been proven for a particular kind of work.
We reject the all-or-nothing bargain that local AI usually offers. People should not have to leave the agent they trust, choose runtimes and quantizations, rebuild their tools, or make a smaller model responsible for an entire workflow. The supervising agent should keep the goal, permissions, tools, and final judgment. Local execution should take on a bounded mission when it fits, work within the authority it was given, and return the result for review with the limits of the run made visible. The winning system is not local instead of cloud. It is a wiser division of labor between them.
The hard problem is not running another model. It is knowing when to delegate, what to ask, and how much to trust the answer. That knowledge should not become another burden for the user. Forjal makes local capacity legible to the supervising agent: what a deployment does well, where it falls short, what the evidence says, and which path this computer can actually run. The agent chooses. Forjal prepares and runs that exact choice without silent substitution. Content sent to the local run stays on the computer by default. Local AI stays inside the workflow people already use, until one day it no longer feels like a separate category at all.