For engineers building multi-step AI workflows

Discover what your AI workflow needs to work reliably.

AQVEN is a Python framework and a local Studio for building reliable LLM workflows. Your coding agent, Claude Code or Codex, runs the experiments, and you see the evidence.

uv tool install aqven
AQVEN Studio showing the graph of a multi-step support workflow
AlphaSource availableAQVEN LicensePythonRuns locallyNo account needed
The research loop

Your agent does the work. You direct the investigation.

  1. YouSet the question.What to find out, what counts as success, and the budget.
  2. AgentDo the experimental work.Prepares variations, runs the experiments, inspects the results, proposes the next step.
  3. AQVENKeep the evidence.Runs the workflow, records every run, applies your checks and computes verdicts.
  4. YouDecide.What to change, what to confirm, and what still needs investigation.
  1. 1

    Build it and run it.

    Agent

    Flows, nodes and prompts as files. Every run leaves a full trace.

  2. 2

    Bet on the riskiest failure.

    YouAgent

    You agree the failure modes. The hypothesis is written before any numbers.

  3. 3

    Explore on working cases.

    AgentAQVEN

    One change at a time. 95% intervals, not one lucky run.

  4. 4

    Confirm on held-out cases.

    AQVEN

    Cases kept apart while exploring. Confirmed, refuted or inconclusive.

  5. 5

    Keep the finding.

    AQVENAgent

    The verdict goes into FINDINGS.md. The failing case is tagged regression.

next round, on fresh cases
Caught by an experiment, not by a customer.
triagegemini-2.5-flash-liteoutput.retries: 2
types/records/observation.yaml · value: maxLength 200
try 1observations[].value over 200 charsrefused
try 2observations[].value over 200 charsrefused
try 3observations[].value over 200 charsrefused

MODEL_RETRIES_EXHAUSTED: counted as schema_invalid, a failed attempt

Triage wrote past its 200-character limit three tries in a row. The engine refused each output, and the agent’s next experiment measures how often it happens.

Evidence

Built for inspection, not blind trust.

AQVEN shows whether a result is promising, confirmed within its scope, inconclusive or invalid.

Check it before it runs.
llmclassify_intent
switchroute_to_queue

E_REF_MISSING: route_to_queue expects category, classify_intent never produces it

aqven check finds broken connections offline, before a run costs a token. It checks the wiring, not whether the answers are right.

Review a change like code.
classify_intent.prompt.md+1-1
Classify the ticket into one of:
-billing, technical, other.
+billing, technical, refund, other.

A prompt edit shows up in your diff. Your teammate reviews it in a pull request.

See where an answer went wrong.
1
collect_orders3 orders read from queue
2
classify_intentcategory: refund, confidence: 0.41
3
route_to_queuesent to: escalations

Every run keeps its full event log. Follow an "almost right" answer back to the step that produced it.

See what each step costs.
collect_orders$0.00 · 12ms
classify_intent$0.02 · 840ms
route_to_queue$0.00 · 3ms

Dollar cost and latency for every step of every run. A call nothing can price counts as unknown, not as free.

Compared with what you use

Traces show. Evals score. AQVEN tells you what to change.

AQVEN compared with Langfuse, LangSmith, promptfoo, DeepEval and n8n
AQVENLangfuseLangSmithpromptfooDeepEvaln8n
Trace of every step of a runYesYesYesYesYesYes
Experiments on your own casesYesYesYesYesYesPartlymetrics on paid plans
Tells you whether a difference is realYesNoNoNoPartlyConfident AI, paidNo
Confirms a change on held-out casesYesNoNoNoNoNo
The workflow as files in your repoYesNoNoNoNoPartlyexport; git on paid plans
Checks the whole workflow before a runYesNoNoNoNoPartlystructure and config
Hosted workspace for a teamNoYesYesPartlyEnterpriseYesYes

Yes Partly, or on a paid plan No, or not documentedChecked against each product’s documentation on 28 September 2026.

Keep your models. Keep your code.

AQVEN runs on your machine, next to your application, with the workflow in your own repo.

flows/route_ticket/nodes/collect_orders/collect_orders.node.yaml
kind: "Node"
node: "code"
description: "Reads unresolved tickets from the queue"
run: "collect_orders"
in:
- name: "queue_id"
type: "QueueId"
from: "$input.queue_id"
out:
- name: "tickets"
type: "Ticket[]"
flows/route_ticket/nodes/collect_orders/collect_orders.py
from route_ticket.types import QueueId, CollectOrdersOut
 
def collect_orders(queue_id: QueueId) -> CollectOrdersOut:
tickets = fetch_open_tickets(queue_id)
return CollectOrdersOut(tickets=tickets)
Source available.

Read the code, run it and change it: your workflows stay yours, whatever happens to us.

lastonoga/AQVEN
AQVEN/
├── packages/
│ ├── aqven/ the engine
│ └── aqven-llm/ model providers
├── apps/
│ └── studio/ the UI
├── LICENSE AQVEN License 1.0.0
└── README.md
$ git clone https://github.com/lastonoga/AQVEN
$ cd AQVEN && uv sync
View on GitHub

Bring your own key. Works with these and the rest of the provider catalog.

OpenAIAnthropicGoogle GeminiOpenRouterMistral AIDeepSeek

Frequently asked questions.

Is AQVEN open source?

No. It is source-available under the AQVEN License 1.0.0: you can read, run and change the code, and commercial redistribution is restricted.

Is AQVEN free?

Yes. You can use it for any purpose, including in production and in commercial products you build with it. The license only rules out competing products and selling AQVEN itself.

Is AQVEN ready for production?

AQVEN is in alpha: expect changes between releases. Try it on a workflow you can afford to change.

Which Python version does AQVEN need?

Python 3.14. uv fetches it for you when you install AQVEN.

What models and providers does AQVEN support?

OpenAI, Anthropic, Google Gemini, OpenRouter, Mistral, DeepSeek and more. Swap a model per agent without touching the rest.

Do I need an account or a cloud service?

Not for AQVEN: it runs on your machine, nothing hosted by us. Model calls use your own provider keys.

Which coding agents does AQVEN work with?

Claude Code and Codex, from Studio's chat or from your own terminal, and any other MCP client, such as Cursor, through AQVEN's MCP server.

Can my coding agent actually edit these workflows?

Yes. Every flow, node and prompt is a plain file your agent can read, edit and check.

Can my coding agent run experiments?

Yes, through MCP tools or the aqven series command, and AQVEN computes the statistics. Near your spend cap, a series waits for a person to continue it in Studio.

Does AQVEN replace my application, or run alongside it?

Alongside it, with your providers and your tools. The workflow itself moves into AQVEN files in your repo: AQVEN doesn't trace one you already have.

What happens to my workflows if AQVEN goes away?

They stay yours. Every flow, node, prompt and dataset is a file in your own repo.

Start with one question.

Run a small experiment. Inspect the evidence. Decide what comes next.