AI + Web3 Integration · 5 min read ·

On-Chain AI Agents: Architecture & Hard Tradeoffs

A practical guide to on-chain AI agent architectures—what to keep on-chain vs off-chain, key tradeoffs, and patterns that actually ship.

On-chain AI agents are the natural collision point between two powerful ideas: autonomous software (agents) and credibly neutral execution (blockchains). Done well, they let you build systems that can decide and act with transparent rules, auditable state, and enforceable permissions—without trusting a single operator.

Done poorly, they become expensive, slow, and insecure “AI theater” glued to a smart contract.

This post breaks down realistic architectures for on-chain AI agents, what belongs on-chain, what shouldn’t, and the tradeoffs founders and developers need to confront early.

What “on-chain AI agent” actually means

An AI agent, at minimum, has:

  1. State: memory, goals, policies, preferences.
  2. Perception: inputs from the world (market data, user requests, events).
  3. Reasoning: an LLM or model that proposes actions.
  4. Actuation: ability to execute transactions—swaps, votes, transfers, calls.

“On-chain AI agent” does not usually mean running an LLM inference inside the EVM (you can, but you probably shouldn’t). In practice, it means:

  • The agent’s authority, constraints, and audit trail live on-chain.
  • The agent’s reasoning loop is often off-chain.
  • The agent’s actions are enforced by smart contracts.

The core question is: what do you need the chain to guarantee? Everything else should be optimized for cost and speed.

Reference architecture: the 5-layer stack

Most shippable systems look like this:

1) On-chain policy + permissions (the “braincase”)

Smart contracts define what the agent is allowed to do:

  • Allowed protocols / function selectors (e.g., only UniswapV3 swap, only Aave repay)
  • Risk limits (max position size, max slippage, max daily spend)
  • Role-based access (who can pause, who can update policy)
  • Time locks and rate limits

This is where you get real value: even if the agent operator is compromised, the contracts constrain damage.

2) On-chain state + accounting (the “memory ledger”)

Keep on-chain what must be auditable or composable:

  • Nonces, task IDs, and execution receipts
  • Portfolio state (holdings already on-chain anyway)
  • Agent configuration hashes (references to off-chain prompts/policies)
  • Performance metrics if you want verifiable reporting

Avoid dumping verbose “chat logs” on-chain. Store commitments (hashes), not paragraphs.

3) Off-chain inference + planning (the “reasoner”)

LLMs are expensive and non-deterministic. Put them off-chain:

  • Read chain state, events, and off-chain signals
  • Generate a plan: one or more contract calls with parameters
  • Optionally run simulation (forked chain, Tenderly/Foundry) before execution

This layer is also where you can swap models as capabilities change—without redeploying contracts.

4) Oracles + attestation (the “truth bridge”)

Agents often need facts not natively on-chain:

  • Prices, volatility, news, social signals
  • Identity / reputation checks
  • “Did X happen?” (e.g., shipment delivered, game result)

Options range from Chainlink-style feeds, to optimistic oracles, to custom committees, to TEEs. Each choice changes the trust model.

5) Execution + settlement (the “hands”)

Execution can be:

  • Direct: agent key signs and sends txs
  • Smart account: ERC-4337 account with session keys and spend controls
  • Keeper network: decentralized bots compete to execute signed intents

The chain is the final arbiter: calls either meet policy constraints or revert.

Three common design patterns (and when to use them)

Pattern A: “Guardrailed operator” (fastest to ship)

  • Off-chain agent proposes actions
  • On-chain contract validates constraints
  • Operator/relayer submits tx

Pros: simple, cheap, great for MVPs.

Cons: availability depends on operator; censorship risk; limited decentralization.

Example: a treasury rebalancer bot that can only swap within a whitelist and daily budget.

Pattern B: “Intent + decentralized execution” (better liveness)

  • Agent signs an intent (EIP-712 message) describing desired outcome
  • Anyone can execute it on-chain for a fee
  • Contract verifies signature + constraints

Pros: better liveness and censorship-resistance; competitive execution; easy auditing.

Cons: more complexity (replay protection, partial fills, MEV considerations).

Example: an agent posts “swap up to 100k USDC for ETH if price < X, max slippage Y”; searchers execute when profitable.

Pattern C: “On-chain verified computation” (maximum integrity, maximum cost)

  • Inference or critical computations are proven via ZK/validity proofs, or verified via TEE attestations
  • Chain verifies proof/attestation before accepting action

Pros: strongest guarantees; less trust in agent operator.

Cons: expensive and hard; limited model flexibility; tooling immature.

Example: a credit/risk scoring step proven by ZK (or attested), used to approve an on-chain loan limit.

The real tradeoffs (the ones that bite later)

1) Determinism vs intelligence

Blockchains want deterministic execution; modern AI is probabilistic.

If you need reproducible outputs, you must either:

  • Make the on-chain part purely rule-based (recommended), or
  • Use verifiable computation/attestations to “pin” what the model did

Most teams should accept that model reasoning stays off-chain and focus on on-chain constraints.

2) Cost and latency vs transparency

Putting more on-chain increases auditability—but quickly becomes unusable.

Practical guidance:

  • Put policies, limits, and approvals on-chain.
  • Put prompts, chain-of-thought, intermediate reasoning off-chain.
  • Anchor with hashes on-chain if you need tamper evidence.

3) Oracle trust is your weakest link

If your agent uses off-chain signals, your security is now the oracle’s security.

Mitigations:

  • Use multiple sources and medianization.
  • Prefer optimistic designs where false data can be challenged.
  • Bound impact: even perfect price feeds shouldn’t let the agent drain the treasury.

4) Key management and blast radius

Agents need keys, but keys get compromised.

Use:

  • Smart accounts with session keys
  • Spend limits and function-level allowlists
  • Emergency pause and time-locked policy upgrades

A good on-chain agent design assumes compromise and limits damage.

5) MEV and adversarial environments

Agents that broadcast predictable actions get sandwiched.

Countermeasures:

  • Private orderflow (e.g., Flashbots / private RPC)
  • Intent-based execution with competition
  • Slippage ceilings and TWAP-style execution

If your agent trades on public mempools without protection, it’s donating to MEV searchers.

Implementation checklist (what we recommend in practice)

  1. Define the trust boundary: what must be guaranteed by the chain?
  2. Write the policy contract first: allowlists, limits, rate controls, pausing.
  3. Use intents instead of raw tx submission when liveness matters.
  4. Simulate before execution: fork tests + revert reason capture.
  5. Log receipts on-chain: task ID, intent hash, outcome.
  6. Design for upgrades safely: immutable core + upgradeable policy parameters via timelock.
  7. Assume hostile conditions: MEV, oracle manipulation, key compromise.

Conclusion: keep intelligence off-chain, keep authority on-chain

The winning architecture for on-chain AI agents is not “LLMs on Ethereum.” It’s verifiable authority on-chain paired with flexible reasoning off-chain.

Put guardrails, accounting, and enforcement in smart contracts. Use off-chain models for planning and iteration. If you truly need stronger integrity, selectively add attestations or proofs—but treat them as scalpel tools, not a default.

On-chain agents are less about making AI trustless and more about making AI accountable. That’s the trade—and the opportunity—for AI + Web3 integration.