AI + Web3 Integration · 5 min read ·

On-Chain AI Agents: Architecture and Tradeoffs

A practical guide to on-chain AI agent architectures, where to place compute and state, and the key security, cost, and UX tradeoffs.

On-Chain AI Agents: Architecture and Tradeoffs

On-chain AI agents are software agents that can observe state, decide, and act through blockchain transactions—often managing funds, interacting with DeFi protocols, and coordinating with other agents. The hard part isn’t the “AI” in isolation. It’s designing a system that is verifiable enough to trust, cheap enough to operate, and fast enough to be useful.

If you’re building an AI + Web3 product, you’ll quickly discover an uncomfortable truth: you can’t put modern model inference on-chain without major compromises. So most “on-chain agents” are really hybrid systems, where some parts are on-chain (state, policies, permissions, attestations), and the heavy compute runs off-chain.

Below is a concrete architecture map and the tradeoffs that matter.

What “on-chain” can realistically mean

There are three broad interpretations:

  1. On-chain state, off-chain intelligence: The chain is the source of truth for agent identity, permissions, treasury, memory pointers, and action history. Decisions are computed off-chain.
  2. On-chain verification of off-chain intelligence: The model runs off-chain, but the chain verifies something about the output (proofs, signatures, stake/slash, or optimistic challenges).
  3. On-chain inference: The model (or a tiny model) runs inside the VM. This is rare today outside of toy models due to gas and latency constraints.

The best architecture depends on your product’s trust assumptions: do users need cryptographic guarantees, or is “auditable with strong incentives” sufficient?

Reference architecture: the agent as a pipeline

A pragmatic on-chain agent system usually splits into five layers:

  1. Identity + permissions (on-chain)

    • Agent registry (address, metadata, version)
    • Roles (owner, operator, guardian)
    • Capability allowlist (which contracts/methods the agent may call)
  2. Policy + constraints (on-chain)

    • Spending limits (per tx/day, per asset)
    • Risk rules (max slippage, min liquidity, max leverage)
    • Emergency controls (pause, withdraw-to-safe, rotate keys)
  3. State + memory (hybrid)

    • On-chain: critical state (positions, budgets, last action nonce, commitments)
    • Off-chain: long-term memory, embeddings, logs, and private context
    • Pointers/commitments: store hashes or content-addressed links (e.g., IPFS/Arweave) to make off-chain memory auditable
  4. Compute + decisioning (off-chain)

    • Data ingestion (RPC, subgraphs, CEX data, news)
    • Model inference (LLM + tools, or a smaller policy model)
    • Simulation (portfolio backtests, MEV/sandwich risk checks)
  5. Execution (on-chain)

    • A transaction builder (EIP-712 typed data, calldata generation)
    • An execution contract (smart account / module / vault)
    • Monitoring + alerting (reorgs, partial fills, failed calls)

This decomposition lets you decide what must be trustless (on-chain) versus what can be trusted-but-audited (off-chain).

Three core design patterns (and when to use them)

1) Smart account agent (practical default)

Pattern: The agent controls a smart account (ERC-4337-style) or vault with strict on-chain policies. Off-chain software proposes actions; on-chain modules enforce constraints.

Why it works: You get strong safety properties without pretending inference is verifiable.

Example: A treasury rebalancer that can only:

  • swap on specific routers,
  • within max slippage,
  • up to a daily budget,
  • and only into an allowlisted asset set.

This is how many production “agents” should start: treat the model as an untrusted recommender and the contract as the enforcer.

2) Attested decisions (good for coordination)

Pattern: Off-chain inference outputs a signed decision (or a batch plan). Validators, a DAO committee, or a decentralized attestation network co-signs it. The chain checks signatures before execution.

Why it works: It’s cheaper than ZK, more decentralized than a single server, and aligns incentives. It also creates a clean audit trail: “who endorsed this action?”

Tradeoff: You’re trusting signers not to collude. This is often acceptable for mid-value actions or where governance already exists.

3) Proved computation (high assurance, high complexity)

Pattern: The agent runs off-chain but submits a proof (ZK/validity proof) that “given inputs X, the model produced output Y.”

Why it matters: This is the closest you get to trustless AI.

Reality check: Proving large transformer inference is still expensive and complex. In practice, teams prove simpler things: risk checks, rule-based policies, or small models. If you need proofs, start by proving constraints and invariants, not full LLM reasoning.

Key tradeoffs you can’t ignore

Cost vs. verifiability

  • Fully on-chain inference: maximal verifiability, brutal gas costs, limited model size.
  • Off-chain inference with on-chain constraints: cheap and robust, but you trust the operator for “intelligence.”
  • Proofs/attestations: a middle ground with overhead.

Opinionated take: Most founders should prioritize constraint enforcement over inference verification. Users care more about “it can’t rug me” than “its chain-of-thought was proven.”

Latency vs. competitiveness (especially in DeFi)

An agent that reacts in 2–10 seconds might be fine for DCA or rebalancing, but it will lose in fast markets to MEV and professional market makers.

Mitigations:

  • Prefer RFQ/intent-based execution (where available)
  • Use private transaction relays / MEV-aware routing
  • Simulate slippage and sandwich risk before signing

Determinism vs. model flexibility

Blockchains like deterministic execution; LLMs are probabilistic.

Practical approach:

  • Use LLMs for planning and tool selection
  • Use deterministic modules for quoting, checks, and execution
  • Log prompts, tool outputs, and final calldata hashes for auditability

Security model: the agent is a hot wallet unless you design otherwise

If your agent can sign transactions, you’ve created a target. Common guardrails:

  • Smart accounts with spending limits and method allowlists
  • Multisig or threshold approvals above a value threshold
  • Time delays for high-risk actions
  • Circuit breakers (pause + withdraw)
  • Key rotation and incident playbooks

A good rule: treat the LLM runtime as compromised. Your on-chain policy should still prevent catastrophic loss.

Data availability and “memory” integrity

Agents rely on off-chain context: positions, prices, historical actions, and user preferences.

If you don’t commit to what the agent saw, disputes become impossible to resolve. Lightweight integrity patterns:

  • Store a hash of the input bundle used for each decision
  • Store the action plan hash (and optionally a pointer to full logs)
  • Use content addressing (IPFS/Arweave) for reproducible traces

A concrete “starter” blueprint

If you want something production-grade without over-engineering:

  1. Smart account/vault contract with:
    • allowlisted targets and function selectors
    • max slippage + oracle sanity checks
    • per-epoch spend limits
    • pause and guardian withdrawal
  2. Off-chain agent service that:
    • reads on-chain state and prices
    • proposes trades with simulation
    • emits signed EIP-712 intents
  3. Executor that:
    • submits transactions privately when needed
    • monitors reverts, reorgs, and drift
  4. Audit trail:
    • store decision input hash + calldata hash on-chain
    • store full logs off-chain, content-addressed

This gives users credible safety and transparency today, without betting your roadmap on bleeding-edge ZK inference.

Conclusion

On-chain AI agents are less about putting an LLM on Ethereum and more about splitting responsibilities: keep identity, permissions, constraints, and critical state on-chain; push heavy compute off-chain; and add verification where it delivers real user value.

The winning systems will be the ones that treat AI as a powerful but fallible component, surrounded by deterministic guardrails. Start with enforceable policies and auditable traces. Add attestations or proofs only when the economics justify the complexity. In AI + Web3, trust is a product feature—and architecture is how you ship it.