Bounded work
Every task carries its own contract: scope, non-goals, a validation plan, and verified completion. Agents can't wander, and “done” means proven done — not asserted done.
Agent Logic
We design systems for serious multi-agent work: bounded execution, reviewable reasoning surfaces, and operational discipline for agent-based tasks.
Company
Agent Logic builds the control plane for AI agents: define the goal, bind the execution, sign the evidence.
Every task carries its own contract: scope, non-goals, a validation plan, and verified completion. Agents can't wander, and “done” means proven done — not asserted done.
Plans, decisions, reviews, and outputs become durable, replayable records — evidence surfaces, not chat residue. Any run can be reconstructed and inspected long after it happened.
Human review stays central while agents take on increasingly complex flows. Approval gates keep authority where it belongs as throughput scales — oversight by design, not by exception.
Every capability an agent gains is also an attack surface — so
our security doctrine gives every important capability an adversarial
counterpart. Under CAV, governed red-team agents hunt for exploits in
the systems blue-team agents defend, and no vulnerability counts as
fixed until its replay bundle fails to reproduce.
Read
about CAV.
See the result of a real CAV security
run
Products
Today, agent systems live in scattered prompts and glue code — impossible to inspect, version, or audit. The Agent Design Language (ADL) defines them explicitly: agents, tools, permissions, workflows, and validation rules, declared in one specification a human can read and a runtime can enforce.
The runtime executes what the language declares. It enforces policy, gates tool use, and records every prompt, call, and decision as a replayable record — deterministic control over non-deterministic models. This is the control plane for enterprise agent systems.
Our first product, and the platform's proving ground: AI-native software engineering built on ADL. Agent teams deliver code through the Cognitive SDLC — every change bound to an issue, validated, and reviewed from four independent perspectives: correctness, security, adversarial, and policy compliance. The runtime keeps the evidence; your engineers keep authority. In development now, with early access through the design-partner program. Learn about CodeFriend - launching soon!
Research
You can't govern what you can't verify. Today, agent behavior is folklore — benchmarks, anecdotes, and vibes. Our research builds the missing substrate: formal ways to plan, validate, trace, and bound what agents do, so behavior becomes something teams can measure and review, not just watch.
The formal foundations are published.
Our Gödel agents run the loop most labs only theorize about: observe their own performance, hypothesize improvements, test them, and adopt what survives. The Gödel–Hadamard–Bayes algorithm makes self-improvement a governed, replayable process — every revision recorded in the reasoning graph, every adoption inside the same gated runtime as everything else. Self-improvement as an engineering discipline, not an article of faith.
Our agents record what they believed and why — claims, evidence, criticisms, hypotheses, and revisions in a graph you can interrogate for contradictions, unsupported conclusions, and unanswered objections. And the graph isn't passive: when it shows tension, the loop runtime generates hypotheses, runs discriminating tests, and revises beliefs. The scientific method, running as software.
Alignment isn't only a model property — it's an architecture property. Every consequential action an agent proposes passes through the Freedom Gate: a policy-and-evidence checkpoint that can allow, refuse, or route to a human. Refusals are recorded with reasons, approvals leave evidence, and operator authority is enforced by the runtime — not requested politely in a system prompt.
Recent publications
We propose a mathematical framework for measuring general intelligence across biological and artificial systems. Intelligence is treated not as an abstract property, but as a measurable quantity: the efficiency with which a system discovers and exploits compact, task-relevant representations while operating within explicit resource constraints. Building on state-space compression (SSC), we develop a formal model of tasks, admissible representations, and resource-aware optimization, leading to Cognitive Compression Cost (CCC)—the minimum achievable cost of solving a task under a declared modeling policy.
Published June 2026 · DOI: 10.5281/zenodo.20738007 ·
Read the paperEvery SDLC ever written assumes humans write the code. C-SDLC is the first lifecycle designed for AI-centric production: it shows how parallel agent teams can approach the Amdahl's Law limit of development speed — and what breaks, and what becomes possible, when the marginal cost of software falls toward zero. For enterprises, it's a blueprint for multiplying SDLC throughput without surrendering quality, review, or control. CodeFriend is where this lifecycle becomes product.
Preprint in review
Tool calls are where agents touch the real world — and every provider defines them differently. UTS is a
proposed open standard for provider-neutral tool interfaces: the same tool runs on any runtime, every call
can be recorded and replayed, and side effects are declared and governed rather than improvised. It's the
interoperability layer a governed control plane needs underneath it.
Read about UTS 1.1
->
Contact
We are speaking with teams working on high-trust agent workflows, multi-agent delivery, and reviewable AI operations.
hello@agent-logic.ai