Browse by type

A programming language and framework for AI orchestration

Running an agent puts one model in charge of the whole job, and you spend time and money waiting for the model to do the task properly.
In Weft, you break down any task to a graph of scoped nodes. Every node (LLM, agents, human in the loop, API, database, ...) gets exactly the context it needs to run, and the graph takes on the task of coordinating each component properly without any nodes having to spend compute on it.
The compiler checks every connection before the run to make sure your orchestration will run properly, and every value that passes between nodes is recorded in a journal you can watch afterwards (unless its trigger turns recorded off).
|
|
**You never have to learn this language.** Tangle is our AI assistant: it knows everything about the language and the framework, and guides you every step of the way. Weft was designed from the ground up to be consumed and produced by AI assistants, and guided by humans. Tangle loads automatically when you open a weft project, and it works with all major assistants: Claude Code, Cline, Codex, Cursor, Devin Desktop, Gemini CLI, GitHub Copilot, Junie, Kilo Code, and OpenCode. |
recorded off), so you can open any of them and see what happened, or is happening, interactively through the graph.pg = PostgresDatabase and the program gets a Postgres of its own. The node says which container it needs; the runtime starts that container with a disk that survives restarts, and the rest of the program reaches it through one output, pg.access. Anything that runs in a container can be an infrastructure node; if you want to write your own, go and read infrastructure nodes.If you scroll a bit on social media, after a few minutes you should have seen a bunch of AI gurus telling you that "you shouldn’t be prompting agents anymore" and instead be designing loops.
What info is missing is what tool you should use for those loops.
Today, making one means building everything around it. An agent framework to express the loop, a durable engine so a wait survives a restart, an observability platform to see it, a policy layer to bound it, a canvas you cannot diff. And even then nobody hands you: the API clients, the database, the queue, the container, the credentials, the form your person actually answers, ... A stack of products, a pile of glue, and the loop you designed lives in the same messy architecture.
Weft is one artifact for all of it. The loop, the API it calls, the database it writes, the container it needs, the person it asks: each are fast to produce and composable nodes.
whatsapp = BaileyBridge
ask = BaileyReceive { bridge: whatsapp.bridge }
draft = LlmInference -> (answer: String, sensitivity: String) {
parseJson: true
prompt: ask.content
provider: OpenRouterProvider { model: "z-ai/glm-5.3" }.provider
params: LlmParams { systemPrompt: @file("prompts/support.md"), temperature: 0.75 }.params
}
# Exactly one of these two says yes, the other stays quiet
route = Switch {
value: draft.sensitivity
cases: [
{ "kind": "equals", "value": "high", "port": "needsAPerson" },
{ "kind": "otherwise", "port": "goAhead" }
]
}
review = HumanQuery {
_should_flow: route.needsAPerson
title: "Send this answer?"
fields: [
{ "kind": "display", "key": "from" },
{ "kind": "display", "key": "question" },
{ "kind": "display", "key": "answer" },
{ "kind": "approve_reject", "key": "send" }
]
from: ask.pushName
question: ask.content
answer: draft.answer
}
allowed = FirstInOrder {
approved: review.send_approved
automatic: route.goAhead
}
reply = BaileySend {
_should_flow: allowed.value
bridge: whatsapp.bridge
to: ask.chatId
message: draft.answer
}
https://github.com/user-attachments/assets/3029cd34-25a8-43ec-b396-e3b37cd0be9c
This is a complete runnable program: a support bot on WhatsApp. A customer messages in, a model writes an answer and rates how sensitive the question was, and if the question is sensitive the program waits for a person to approve the answer before sending it.
The first line is the WhatsApp integration. It is only one line because BaileyBridge is an infrastructure node: it ships with its own container and manages its own lifecycle. After you start the infrastructure, the bridge shows a QR code in the graph view. Scan it with whatsapp to connect the bridge. You can check the source code of BaileyBridge in the catalog/ if you want to see how it works.
Then ask is a trigger: when a message arrives, it starts an execution with the payload. draft then calls an LLM and drafts a reply. The system prompt asks for JSON. parseJson: true parses the reply, attempts to repair malformed JSON, and pulls out the declared outputs: the answer and the sensitivity rating. The prompt lives in its own file, and the model is pinned in the source, here z-ai/glm-5.3.
After that there is a Switch. It looks at draft.sensitivity and opens exactly one of its two ports: needsAPerson if the rating is "high", goAhead otherwise. review reads that port through _should_flow, so the person is only asked when the answer was sensitive.
When the review fires, this execution waits for a human to review the answer through the integrated Weft browser extension. The question and execution state are saved. Once the worker has no other work, it exits; the waiting execution needs no process of its own. When the human replies, the runtime restores the recorded state and continues. The WhatsApp bridge and shared runtime services stay up. For what survives a restart, read the journal.
If the user refuses, send_rejected fires and send_approved closes. That propagates through the rest of the program and stops the execution. If the user approves, send_approved fires and send_rejected closes. (The two are auto-inferred ports from the field of kind approve_reject.)
FirstInOrder emits the first of its inputs, in written order, that delivered a value: review.send_approved or route.goAhead. reply uses that as its _should_flow.
The whole thing is a runnable project in examples/whatsapp-support-bot/.
You need Docker, and
Rust if the installer has to build the CLI itself (it
tells you when). The installer sets up weft's runtime on your machine, the
weft command and the VS Code extension. The extension also installs from
the store in VS Code
forks such as Devin Desktop. We are working on a CLI-only version, support for other IDEs and a cloud hosted version. If the extension does not install
automatically, follow the manual instructions the installer prints.
git clone https://github.com/WeaveMindAI/weft.git
cd weft
./setup.sh
weft new hello --assistant claude-code --remember
cd hello
weft run
Open hello/main.weft in VS Code to see the program as a graph. Open the
same project in your coding assistant and tell Tangle what you want to build.
Claude Code is the example here; you pick your assistant with --assistant, and
--remember makes it the default, so your next weft new needs no flag.
If the installer reports a missing tool, it prints where to get it. If your
shell cannot find weft, use the PATH line it printed. For installation
options and help, go to Install.
The installer picks the build for you: an unchanged checkout that matches a published build gets the prebuilt CLI and editor, and anything else is built locally with the reason printed. If it asks for extension build tools, follow Contributing.
For human questions and approvals, install Weft tasks from Firefox Add-ons or the Chrome Web Store.
To connect Weft tasks to weft, read the human-step walkthrough.
weft run --seed reuses compatible completed work so you are not paying
for the same steps again. Once a run comes out right, weft freeze <name>
keeps its starting parameters and the outputs you accepted, weft run <name>
replays them against the code you have now, and weft diff shows you or
Tangle what changed. Once the whole program works on one example, try
another and keep the earlier examples passing. Read
Sequential Diffusion Programming
for that process, and Versions, seeded runs and frozen examples
for the verbs.weft describe-nodes --list
lists the available nodes. If the one you need is missing, start with
Writing a node.If you wanted to build a rocket, you wouldn't look for one unbelievably clever person and tell them to get on with it. You'd put smart people together, give them tools and resources, and think hard about how they should work together. Getting AI to work in production deserves the same attention. A more capable model helps, of course, but we are betting that improving the thing around it matters just as much.
And there is something fundamental about LLMs that you must keep in mind: what we call “the AI assistant” isn't the model, even though most people talk as if they are the same thing. The model predicts probabilities for the next token, then a sampling algorithm chooses which token actually happens. That token goes back into the context, the model predicts again, and on it goes. Out of this repeated prediction and sampling c
browse all types & interfaces →