Cloud and local AI, working as one.
Forjal is an applied AI lab. We research and build the systems that let the AI you use in the cloud work with the computer you already own.
Read the manifestoComputing has done this before.
Work moved from the mainframe to the personal computer when personal machines became efficient enough for the job, not when they beat the mainframe. AI is at the same turn.
more intelligence per watt from local AI, 2023 to 2025
Stanford
for open models to double their capability per parameter
Nature Machine Intelligence
of new PCs expected to ship with an AI chip in 2026
Counterpoint Research
Your agent has one place to send the work.
Reading a file, running a test, re-reading its own history: every step goes to the same frontier model in the same data center. Not because each step needs it, but because there is nowhere else.
So you pay frontier prices for work that never needed a frontier model. Everything the agent touches leaves your computer. And the machine in front of you, with a chip built for AI, has no part in any of it.
NVIDIA researchers argue that small models are powerful enough, better suited and more economical for many of the calls an agent makes. How much of the work could actually move, and which parts, is still an open question. It is the question our research is built around.
Source: Belcak et al., NVIDIA Research, 2025

Local AI should join the AI you use.
Not replace it.
We are early. The best questions are still open.
Research and product share one loop: every product we ship is an instrument of the research.
All research →Where local AI fits.
Which parts of people’s work belong on their own machine.
The hardware people own.
How to get the most from CPUs, GPUs and NPUs.
Capable local agents.
The architecture and harness that make local models work.

Delegation runs on proof
A visible call is not a delivery, and a correct answer is not a proven one. For a cloud agent to accept a local agent's work, the result has to carry literal, located and sufficient evidence and a record from the executor. That points to a way to split the work: delegate what is cheaper to check than to do.
Read the noteForjal Alpha
Keep your cloud agent. Add a local one. Forjal Alpha adds a local agent to the AI agents you already use.
Install.
Forjal finds a model that fits your machine and prepares it for the accelerator that machine has.
Connect.
Claude Code, Codex, OpenClaw, Hermes Agent, OpenCode or Pi, from the app.
Delegate.
Your agent hands off a task and keeps working.
Only what is worth handing off.
Handing off everything can cost more than it saves. Forjal tells your agent what each local model is good for and what it is not, so the work that goes local is the work that holds up.
Built for the accelerator in your machine.
Where an accelerator lane is not ready, Forjal prepares a GPU or CPU path instead. The choice happens when it prepares the model, and nothing switches during a run.
- Claude Code
- Codex
- OpenClaw
- Hermes Agent
- OpenCode
- Pi
Work with us
Bring it to your team.
Your company already owns the compute. The laptops your team uses carry chips built for AI, and Forjal puts them to work inside the agents your team already uses. We are working with a few teams on what that should look like.
Build with us.
We are a small lab. If you want to work on how cloud and local AI work together, we would like to hear from you.
Make local intelligence a natural part of every AI workflow.
Your hardware, in the loop.




