Skip to content

About

No description, website, or topics provided.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

Agentic Feedback Framework

Errors should not be a wall. They should be a door. A small, language-agnostic protocol for sending structured, machine-readable recovery instructions when an API call fails. Built for autonomous AI agents that interact with HTTP APIs — but useful for any client that wants programmatic error recovery.

License: MIT Node: >=18 Tests: 15/15


Table of contents

  1. The thesis
  2. What this is
  3. Why this exists (the problem)
  4. Why this isn't a memory problem
  5. What this prevents
  6. Why both memory and feedback
  7. Quick start
  8. Quickstart for AI agents
  9. The feedback envelope (the data shape)
  10. Server-side usage
  11. Agent-side usage
  12. The action vocabulary
  13. The registry
  14. Discovery: GET /api/errors
  15. Testing
  16. Adding new error codes
  17. Language support
  18. FAQ
  19. License and contributing

The thesis

An API should not just accept requests and return failures. It should guide autonomous systems toward successful outcomes.

If the server already knows why a request failed and what the correct input should have been, the only sensible thing to do is tell the agent. Every failed tool call the server could have explained wastes time, burns tokens, and increases the chance of the agent hallucinating a solution.

Success responses have been carefully designed for decades. Error responses have not — they were designed for developers reading them in a browser, not for agents reading them in a hot loop with a context budget. As autonomous agents become first-class users of APIs, error responses need to evolve. They should be just as informative, just as consistent, and just as machine-readable as success responses.

This framework is the contract that makes that real: a small, language-agnostic envelope that carries the expected shape, the actual offending values, the recovery hint, the retry decision, and the controlled-vocabulary action — in a single response body that an agent can dispatch on without an LLM call.

The framework is not a memory layer and it does not replace one. It makes a different claim: deterministic recovery should be part of the API contract, not a side effect of an agent's prompt.


What this is

A protocol and a small reference implementation. The protocol says: when your API rejects a request, the response body should be a structured envelope that includes not just the error message, but:

  • What the server expected (in human + LLM readable form)
  • What the server saw (echo of the offending values)
  • A drop-in replacement request body the client can use to retry
  • Whether retrying is even worth it (retryable: true/false)
  • A controlled-vocabulary action the client should take
  • A docs URL for the human reviewer

The reference implementation is ~500 lines of Node.js that provides the registry, the renderer, an Express adapter, and a registry self-test. You can use the reference implementation directly, or re-implement the protocol in any language you want — the wire format is just JSON.


Why this exists (the problem)

When a human hits a 400 in a browser, the error string is usually enough. They read it, look at the form, figure out what went wrong, and try again.

When an AI agent hits a 400, the LLM has to:

  1. Parse a free-text error message (often vague, like "invalid input")
  2. Guess what the server actually wanted
  3. Re-engineer the request, often by hallucinating new parameters
  4. Burn a reasoning cycle to debug the guess
  5. Repeat until it works (or hits a context-window limit)

This is expensive (tokens), unreliable (the LLM hallucinates), and debug-hostile (you can't tell from the agent's logs whether the API was broken or the agent misunderstood the error).

The Agentic Feedback Framework replaces the free-text error with a structured envelope that:

  • Tells the agent the expected shape of the request
  • Tells the agent the exact value the server rejected
  • Provides a drop-in replacement body to retry with
  • Tells the agent whether retrying is even worth it
  • Tells the agent which controlled-vocabulary action to take
  • Provides a docs URL so the human reviewer can read the canonical explanation

The envelope is small (a few hundred bytes), always present on errors, and uses a registry of error codes that's shared between server and client. Every code has a fixed shape, a fixed set of template variables, and a single semantic meaning.

What it looks like

Here's a real exchange. An agent calls an order-pricing endpoint and asks for a price quote with a tight tolerance:

POST /api/orders/quote HTTP/1.1
Content-Type: application/json

{
  "productId": "sku-42",
  "side": "buy",
  "quantity": 600,
  "maxPricePerUnit": 600,
  "customerId": "cust-7821"
}

The server rejects (the price has drifted past the 1% tolerance since the quote was issued). Without the framework, the response would be:

{
  "ok": false,
  "error": "price has drifted past tolerance"
}

The agent now has to guess what "drifted past tolerance" means, what the current price is, and how to fix the request. It will likely fail several times.

With the framework, the response is:

{
  "ok": false,
  "error": "maxPricePerUnit 600 exceeds quoted price * 1.01 (594)",
  "code": "PRICE_DRIFTED",
  "feedback": {
    "expected": "maxPricePerUnit must be no more than 606 (1% above the current quote of 600)",
    "current": {
      "maxPricePerUnit": 600,
      "currentQuote": 600
    },
    "suggestion": {
      "maxPricePerUnit": 606
    },
    "retryable": true,
    "action": "retry_with_higher_maxPricePerUnit",
    "docs_url": "https://your-api.com/errors/PRICE_DRIFTED"
  }
}

The agent's client code can now do:

if (body.feedback.retryable && body.feedback.suggestion) {
  const retried = { ...req.body, ...body.feedback.suggestion };
  return fetch(req.path, { method: 'POST', body: JSON.stringify(retried) });
}

And it works. No LLM reasoning, no hallucination, no token waste.


Why this isn't a memory problem

Many people ask: isn't this just what agent memory is for? Store the API's quirks in long-term memory and the agent learns from past failures.

Memory is valuable. But memory alone does not solve the biggest challenge facing autonomous agents: reliable execution in unfamiliar environments.

Three realities make memory insufficient on its own:

  1. Not every agent has persistent memory. Some agents run stateless — every request is the first request. Others lose context through memory compression, summarization, or session restart. Even an agent with perfect memory cannot know how to use every API, tool, or application it encounters for the first time.

  2. Memory does not survive contact with the API. An agent that learned yesterday that maxPricePerUnit must be above 95% of the current quote might, after a memory compaction, "remember" the rule as must be below 95% of the current quote. The next request fails. The memory was wrong, and the agent has no way to know.

  3. Memory cannot reason about server state. Even with perfect recall, an agent cannot deduce that a resource it doesn't know exists is required, or that the retry budget is exhausted, or that the only safe retry is with a specific token it has never seen before. These are facts that exist only on the server side at the moment of the request.

The Agentic Feedback Framework addresses what memory cannot: it makes the API itself tell the agent what it needs to know to recover on the current request, regardless of what the agent remembers from the past. Memory lets an agent learn from yesterday; feedback lets it succeed today.


What this prevents

The framework is a defense against six classes of failure that otherwise happen constantly in agentic API interactions:

1. Silent breakage from vague errors

Without the framework: an agent sees "error": "invalid request" and either gives up or tries 20 random things. From the operator's perspective, the API is "broken" or "the agent is broken." No one knows which, and the failure mode is invisible until you read the agent's reasoning trace.

With the framework: every error has a code and a docs_url. The agent either auto-recovers (because retryable: true + suggestion) or surfaces a clear, actionable error to the operator.

2. Hallucinated parameters

Without the framework: an LLM looks at "maxPricePerUnit too low" and "fixes" it by changing quantity instead — a plausible-sounding but completely wrong guess. Now the order is structured for the wrong quantity, and the user just spent tokens on nonsense.

With the framework: the agent's handler is dispatched by action, not by parsing free text. The action: "retry_with_higher_maxPricePerUnit" handler is hard-coded to raise maxPricePerUnit; it can't hallucinate.

3. Wasted reasoning cycles

Without the framework: a 5-step agent task fails on step 3. The agent re-thinks the whole task from scratch, possibly reaching step 3 with a different (and worse) plan. The user pays for a 2x LLM cost.

With the framework: the agent catches the feedback envelope, applies the suggestion, and continues. One step, one retry, no re-planning.

4. Unsafe retries

Without the framework: an agent decides "this 500 is probably transient, let me retry 5 times." Now you've DDoS'd your own API, or worse, you've broadcast a half-signed transaction 5 times.

With the framework: every error has an explicit retryable flag. If retryable: false, the agent doesn't retry. The action vocabulary also prevents dangerous retry patterns (e.g. report_bug and contact_operator never auto-retry).

5. Operator-on-call for things the API could have answered

Without the framework: an agent gets stuck. It pages the human operator. The operator reads the error, says "you forgot to set direction," and un-sticks the agent. This happens 50 times a day.

With the framework: the API tells the agent exactly what was missing. The agent fixes it. The operator only gets paged for things the API genuinely can't answer (server bugs, missing funds, etc.)

6. The same mistake across sessions

Without the framework: an agent learns that this particular API needs Content-Type: application/json and stores the lesson in memory. After a memory compaction, summarization, or session restart, the rule is gone. The agent makes the same wrong request on its first call to the API in the new session. The mistake repeats, the tokens re-burn, and the user wonders why the agent is "still broken."

With the framework: the recovery is encoded in the response itself. It doesn't matter whether the agent has any prior context — the very first request after a fresh session will get the same structured envelope, with the same suggestion, and the agent will recover in one round-trip. Recovery becomes part of the API, not part of the agent's brain.


Why both memory and feedback

Memory and feedback solve different problems and should not be confused.

Memory Feedback
Timescale Past sessions, yesterday's mistake The current request, right now
Lifetime Persists across sessions Lives in a single HTTP response
Subject to compression? Yes — every compaction is a chance for the rule to be lost, summarized away, or contradicted No — the response is the source of truth at the moment of the call
Can describe server state? No — only what the agent remembers having seen Yes — the server's current view of the request, the resource, the constraints
Can compute a fix? No — memory stores what was, not what should be Yes — the server can return a suggestion object computed from current state
Failure mode Compaction drift, hallucination, "I remember this wrong" Not used, ignored, or poorly mapped to an action

Memory allows an agent to learn from the past. A feedback framework allows an agent to succeed in the present. For the future of agentic ecosystems, both are important. But without deterministic recovery, even the most intelligent models will continue wasting resources trying to solve problems the application already knows how to fix.

The two complement each other:

  • Use memory to avoid re-discovering what an API generally does.
  • Use feedback to recover from what this specific request got wrong, in this specific moment, on this specific server.

A well-designed agent has both. A well-designed API offers both implicitly — feedback as the contract, memory as a client-side optimization that the server doesn't need to know about.


Quick start

Install

npm install agentic-feedback-framework

(You can also just cp src/agentic-feedback.js src/express-feedback.js into your project — there's no compilation step.)

Wire it into an Express app

import express from 'express';
import {
  sendError, sendOk, jsonParseErrorHandler, errorsEndpoint,
} from 'agentic-feedback-framework/express';
import { failWithFeedback, makeErrorWithFeedback } from 'agentic-feedback-framework';

const app = express();
app.use(express.json({ limit: '100kb' }));
app.use(jsonParseErrorHandler);
app.get('/api/errors', errorsEndpoint({ docsBase: 'https://your-api.com/errors' }));

app.post('/api/things', (req, res) => {
  if (!req.body?.name) {
    return sendError(res, 400, 'MISSING_FIELD', 'name is required', {
      field: 'name',
      got: req.body?.name,
      current: { name: req.body?.name },
    });
  }
  // ... do the thing
  sendOk(res, { thing: stored }, 201);
});

Run the example

npm run example:server        # in one terminal
npm run example:agent         # in another
python example/agent_client.py  # or in a third

The example server is a complete, runnable Express app at example/server.js. The example clients (Node and Python) demonstrate how to consume the feedback envelope on the agent side.


Quickstart for AI agents

If you are an LLM-powered agent calling an API that speaks this framework, this is the only section you need to read. Everything below is reference material; the recovery loop fits in ~30 lines.

The principle

When your request fails, do not parse the free-text error string. Read the feedback envelope and execute the action:

action value What you do
retry_with_suggestion Merge feedback.suggestion into your request body, resend.
pick_different_value Same — the server computed a safe replacement (e.g. slug).
batch_smaller Split your batch down to feedback.suggestion.batchSize.
wait_and_retry Sleep feedback.suggestion.retry_after_s seconds, resend.
abort_or_retry Same as wait_and_retry for an in-progress operation.
reauthenticate Re-obtain credentials, resend. Surface docs_url to operator.
fix_input_format Merge feedback.suggestion into the request body, resend.
fix_address_format Merge feedback.suggestion into the request body, resend.
use_different_field Merge feedback.suggestion into the request body, resend.
provide_missing_context Surface feedback.expected to your operator/loop, retry.
wait_for_user_action Stop. Show docs_url to a human.
report_bug Stop. Surface docs_url to the operator.
contact_operator Stop. Surface docs_url to the operator.
no_action_possible Stop. The operation genuinely cannot proceed.

If feedback.retryable is false, do not retry — surface to the operator and stop. The current field always echoes the values the server saw; use it for debug logs.

The reference recovery loop (Node)

async function callWithRecovery(method, path, body) {
  const r = await fetch(`https://api.example.com${path}`, {
    method, headers: { 'content-type': 'application/json' },
    body: body ? JSON.stringify(body) : undefined,
  });
  const json = await r.json();
  if (r.ok || !json.feedback) return json;

  const { action, suggestion, docs_url, expected } = json.feedback;
  switch (action) {
    case 'retry_with_suggestion':
    case 'pick_different_value':
    case 'fix_input_format':
    case 'fix_address_format':
    case 'use_different_field':
      return callWithRecovery(method, path, { ...body, ...suggestion });
    case 'batch_smaller':
      return callWithRecovery(method, path,
        { ...body, items: body.items.slice(0, suggestion.batchSize) });
    case 'wait_and_retry':
    case 'abort_or_retry':
      await new Promise(r => setTimeout(r, (suggestion.retry_after_s || 30) * 1000));
      return callWithRecovery(method, path, body);
    case 'reauthenticate':
    case 'provide_missing_context':
    case 'wait_for_user_action':
    case 'report_bug':
    case 'contact_operator':
    case 'no_action_possible':
      throw new Error(`${action}: ${expected} (${docs_url})`);
  }
}

The reference recovery loop (Python)

import json, time, urllib.request

def call_with_recovery(method, path, body=None):
    data = json.dumps(body).encode() if body else None
    req = urllib.request.Request(
        f"https://api.example.com{path}",
        data=data, method=method,
        headers={"content-type": "application/json"},
    )
    try:
        with urllib.request.urlopen(req, timeout=10) as r:
            return json.loads(r.read())
    except urllib.error.HTTPError as e:
        resp = json.loads(e.read())
        fb = resp.get("feedback") or {}
        action, sug = fb.get("action"), fb.get("suggestion") or {}
        if action in ("retry_with_suggestion", "pick_different_value",
                      "fix_input_format", "fix_address_format",
                      "use_different_field"):
            return call_with_recovery(method, path, {**(body or {}), **sug})
        if action == "batch_smaller":
            return call_with_recovery(method, path,
                {**(body or {}), "items": body["items"][:sug["batchSize"]]})
        if action in ("wait_and_retry", "abort_or_retry"):
            time.sleep(sug.get("retry_after_s", 30))
            return call_with_recovery(method, path, body)
        # Unrecoverable — surface to operator.
        raise RuntimeError(f"{action}: {fb.get('expected')} ({fb.get('docs_url')})")

Discovery: learn the API's vocabulary once

Before your first request, hit the discovery endpoint:

curl https://api.example.com/api/errors

You will get a registry of every error code, every action, and every docs URL. Use it to build your recovery handler table. Cache the result — re-fetch only when the API ships a new version.

Worked example: the loop closed

The agent sends POST /api/notes with {"slug":"welcome"}. The server replies:

{
  "ok": false,
  "code": "ALREADY_EXISTS",
  "error": "a note with slug \"welcome\" already exists; try \"welcome-2\"",
  "feedback": {
    "expected": "a note with value \"welcome\" already exists; pick a different value",
    "current": {},
    "suggestion": { "value": "welcome-2", "field": "slug" },
    "retryable": true,
    "action": "pick_different_value",
    "docs_url": "https://your-api.com/errors/ALREADY_EXISTS"
  }
}

The agent does not parse the error string. It reads feedback.action = pick_different_value, merges feedback.suggestion into the body (slug → welcome-2), and resends. Server replies 201. Total reasoning spent on this round-trip: zero.

Agent-input safety (server-side note for API authors)

This framework also catches the other direction: agents that send malformed requests. If your API serves autonomous agents, use:

  • strictJson({ maxBytes: '100kb' }) instead of express.json()
  • requireJsonContentType (415 if agent forgets Content-Type)
  • requireJsonObject({ minFields: 0 }) (400 if body is null/array/primitive)
  • requireMethod('POST') (405 with suggestion.method if wrong verb)
  • notFoundHandler({ discoveryPath: '/api/errors' }) (404 with hint)
  • errorEnvelope() (top-level catch-all for thrown errors)

Without these, an agent that sends [1,2,3] as a body or GETs a POST-only route gets Express's HTML 404/413/415 wall — and burns hundreds of tokens guessing what went wrong.


The feedback envelope (the data shape)

The envelope is attached to every failed response, at the top level:

type ErrorResponse = {
  ok: false;                  // always false on errors
  error: string;              // short human-readable summary
  code: string;               // machine-readable error code (in FEEDBACK_REGISTRY)
  feedback: FeedbackEnvelope; // the agentic feedback envelope
};

type FeedbackEnvelope = {
  expected: string;           // what the request SHOULD have been
  current: object;            // echo of the offending input value(s)
  suggestion: object | null;  // drop-in replacement request body, or null
  retryable: boolean;         // whether retrying with `suggestion` will likely succeed
  action: ActionString;       // one of FEEDBACK_ACTIONS — a controlled vocabulary
  docs_url: string;           // canonical explanation for the human reviewer
};

And the success response:

type SuccessResponse = {
  ok: true;                   // always true on success
  // ... your data fields
};

The framework doesn't constrain what your success response looks like, other than the ok: true convention. The point is to make the discrimination between success and failure unambiguous to a parser (it can't both be true and false, ever).

expected

A human + LLM readable description of what the request should have been, with template variables filled in. For example:

"maxPricePerUnit must be no more than 606 (1% above the current quote of 600)"

The expected field is what an LLM would have tried to hallucinate from a vague error. By providing it directly, you save the LLM a reasoning cycle AND eliminate hallucination risk.

current

An echo of the offending input value(s). The server saw these specific values; the agent can use this to confirm what the server actually parsed (catches cases where the agent and server disagree on the request body).

suggestion

A drop-in replacement request body the agent can merge into its current request and retry. If null, no automatic fix is possible — the agent should surface the error to the operator.

A typical suggestion might be:

{
  "maxPricePerUnit": 606
}

The agent merges this with the original request and retries. If the retry succeeds, the agent never has to re-plan the task.

retryable

true if applying the suggestion is likely to succeed. false if the agent should give up and surface the error.

Examples:

  • PRICE_DRIFTED: retryable: true (raise the max to the new quote)
  • INVALID_FIELD: retryable: false (the format is wrong; the agent needs human help to know the right format)
  • INSUFFICIENT_FUNDS: retryable: true (the agent or human can fund the account and retry)
  • NOT_FOUND: retryable: false (the resource doesn't exist; retrying with the same id will fail the same way)

action

A controlled-vocabulary string from FEEDBACK_ACTIONS. The complete set is defined in the framework, and adding a new action is a breaking change for agents (it requires a new handler on the agent side). The framework ships with a default set:

retry_with_suggestion
wait_and_retry
abort_or_retry
reauthenticate
pick_different_value
use_different_field
fix_input_format
fix_address_format
batch_smaller
wait_for_user_action
provide_missing_context
report_bug
contact_operator
no_action_possible

docs_url

A URL pointing to the canonical explanation of this error code. The agent doesn't usually follow this link, but the human reviewer of the agent's logs will. Make these pages real and useful.


Server-side usage

Pattern 1: return-style validator (caller already returns)

import { failWithFeedback, FEEDBACK_REGISTRY } from 'agentic-feedback-framework';

function validateOrder(body) {
  if (!body.productId) {
    return failWithFeedback('MISSING_FIELD', 'productId is required', {
      field: 'productId',
      got: body.productId,
      current: { productId: body.productId },
    });
  }
  if (body.maxPricePerUnit < 0) {
    return failWithFeedback('INVALID_FIELD', 'maxPricePerUnit must be >= 0', {
      field: 'maxPricePerUnit',
      reason: 'negative value',
      got: body.maxPricePerUnit,
      current: { maxPricePerUnit: body.maxPricePerUnit },
    });
  }
  return { ok: true };
}

app.post('/api/orders/quote', (req, res) => {
  const v = validateOrder(req.body);
  if (v && v.ok === false) {
    return sendError(res, 400, v.code, v.message, v.feedback_params);
  }
  // ... handle the order quote
});

Pattern 2: throw-style validator (caller has try/catch)

import { makeErrorWithFeedback } from 'agentic-feedback-framework';

async function loadProductOrThrow(productId) {
  const product = await db.findProduct(productId);
  if (!product) {
    throw makeErrorWithFeedback('NOT_FOUND', `product ${productId} not found`, {
      resource: 'product',
      got: productId,
    });
  }
  return product;
}

app.post('/api/orders', async (req, res) => {
  try {
    const product = await loadProductOrThrow(req.body.productId);
    // ... create the order
  } catch (err) {
    // err.code and err.feedback are attached by makeErrorWithFeedback
    if (err.code) {
      return sendError(res, 400, err.code, err.message, err.feedback?.current || {});
    }
    throw err;  // unknown error, let Express handle it
  }
});

Pattern 3: inline error in a route

app.post('/api/things', (req, res) => {
  if (req.body.name.length > 100) {
    return sendError(res, 400, 'INVALID_FIELD', 'name too long', {
      field: 'name',
      reason: 'max 100 chars',
      got: req.body.name,
      current: { name: req.body.name, length: req.body.name.length },
    });
  }
  // ...
});

Pattern 4: discovery endpoint

import { errorsEndpoint } from 'agentic-feedback-framework/express';

app.get('/api/errors', errorsEndpoint({ docsBase: 'https://your-api.com/errors' }));

This serves two routes:

  • GET /api/errors → { docs_base, summary, codes } (full registry)
  • GET /api/errors?code=PRICE_DRIFTED → single code's full feedback envelope, with empty params filled in, plus the list of template parameters

Customization

If you have your own error code names, add them to FEEDBACK_REGISTRY:

import { FEEDBACK_REGISTRY, FEEDBACK_ACTIONS } from 'agentic-feedback-framework';

// Extend the registry (must happen at module load time, before
// validateRegistry() is called)
Object.assign(FEEDBACK_REGISTRY, {
  MY_CUSTOM_CODE: {
    action: FEEDBACK_ACTIONS.FIX_INPUT_FORMAT,
    retryable: false,
    expected_template: 'field {field} has the wrong flavor of wrongness: {got}',
    suggestion_template: null,
    docs_url: 'https://your-api.com/errors/MY_CUSTOM_CODE',
  },
});

For a more permanent extension, fork the framework and add your codes to src/agentic-feedback.js directly.


Agent-side usage

The agent's job is to inspect the feedback envelope and dispatch to a handler based on the action value. The simplest possible client:

async function callApi(method, path, body) {
  const resp = await fetch(path, {
    method,
    headers: { 'content-type': 'application/json' },
    body: body ? JSON.stringify(body) : undefined,
  });
  const data = await resp.json();
  if (data.ok) return data;

  const fb = data.feedback;
  if (!fb) throw new Error(`${data.code}: ${data.error}`);

  if (fb.retryable && fb.suggestion) {
    // Try the server's suggested fix
    return callApi(method, path, { ...body, ...fb.suggestion });
  }

  // Unrecoverable — surface the docs URL
  throw new Error(`${data.code}: ${fb.expected} (see ${fb.docs_url})`);
}

A more thorough client (the one in example/agent-client.js) has a handler for every action value:

const actionHandlers = {
  // Generic — the server's `suggestion` body is a drop-in fix.
  retry_with_suggestion: (req, fb) => retryWith(req, fb.suggestion || {}),

  // Pause + retry — server state will recover on its own.
  wait_and_retry: async (req, fb) => {
    await sleep((fb.suggestion?.retry_after_s || 30) * 1000);
    return retryWith(req, {});
  },

  // Same as wait_and_retry but for an in-progress operation.
  abort_or_retry: async (req, fb) => {
    await sleep((fb.suggestion?.retry_after_s || 30) * 1000);
    return retryWith(req, {});
  },

  // Re-fetch credentials and retry.
  reauthenticate: async (req, fb) => {
    throw new UnrecoverableError('agent must refresh credentials');
  },

  // Server has a different handle / slug / identifier it can suggest.
  pick_different_value: (req, fb) => retryWith(req, fb.suggestion || {}),

  // Wrong field name was used; merge the suggested replacement.
  use_different_field: (req, fb) => retryWith(req, fb.suggestion || {}),

  // Wrong type / format; merge the suggested replacement.
  fix_input_format:    (req, fb) => retryWith(req, fb.suggestion || {}),

  // Address on the wrong network; merge the suggested replacement.
  fix_address_format:  (req, fb) => retryWith(req, fb.suggestion || {}),

  // Server has a max batch size in suggestion.batchSize.
  batch_smaller: (req, fb) =>
    retryWith(req, { ...req.body, items: req.body.items?.slice(0, fb.suggestion?.batchSize) }),

  // Stop — a human needs to take action (fund account, approve, etc.).
  wait_for_user_action: (req, fb) => {
    throw new UnrecoverableError(`a human must take action — see ${fb.docs_url}`);
  },

  // Stop — agent needs to fetch missing context (a parent ID, etc.).
  provide_missing_context: (req, fb) => {
    throw new UnrecoverableError(`agent needs more context — see ${fb.expected} (${fb.docs_url})`);
  },

  // Stop — server-side bug; report to operator.
  report_bug: () => { throw new UnrecoverableError('server bug; report to operator'); },

  // Stop — operator action required.
  contact_operator: () => { throw new UnrecoverableError('operator action required'); },

  // Stop — generic catch-all for unrecoverable errors.
  no_action_possible: (req, fb) => {
    throw new UnrecoverableError(`${fb.expected} (${fb.docs_url})`);
  },
};

A working Python version is in example/agent_client.py.


The action vocabulary

The set of action strings is small and stable. Adding a new action is a breaking change for agents, so the framework only adds new actions when a genuinely new class of recovery becomes common. The complete set, grouped by what the agent should do:

"Just retry with the server's suggestion"

Action When to use
retry_with_suggestion Generic — merge feedback.suggestion into the request and resend.
batch_smaller Server has a suggestion.batchSize; shrink and resend.
pick_different_value The server computed a safe replacement value; merge and resend.
use_different_field The wrong field was used; switch to the suggested field.
fix_input_format A field has the wrong type/format; merge suggestion and resend.
fix_address_format An address is on the wrong network; merge suggestion and resend.

"Pause and retry"

Action When to use
wait_and_retry Server state will recover; sleep and resend.
abort_or_retry Same, but for an in-progress operation that may need abort.

"Reauthenticate and retry"

Action When to use
reauthenticate Credentials are missing or expired; refresh and resend.

"Stop — surface to a human or operator"

Action When to use
wait_for_user_action A human must do something first (fund account, approve, etc.).
provide_missing_context The agent must supply missing context (a parent ID, etc.).
report_bug Server-side invariant violation; report to operator.
contact_operator Operator action is required (config, permissions, etc.).
no_action_possible Generic catch-all for unrecoverable errors.

The registry

Every error code is a key in FEEDBACK_REGISTRY. The shape:

{
  SOME_CODE: {
    action:              'fix_input_format',  // from FEEDBACK_ACTIONS
    retryable:           true,                // boolean
    expected_template:   'field {field} ...', // string with {param} placeholders
    suggestion_template: { field: '...' },    // object or null
    docs_url:            'https://...',       // must be https
  },
}

The framework ships with 10 starter codes (MISSING_FIELD, INVALID_FIELD, RATE_LIMITED, INTERNAL, MISSING_AUTH, INVALID_KEY, FORBIDDEN, NOT_FOUND, ALREADY_EXISTS, STATE_CONFLICT) that cover the 90% case for a generic API. Add more as your application needs them — see "Adding new error codes" below.


Discovery: GET /api/errors

The framework exposes the registry as an HTTP endpoint so agents can introspect the error space at runtime:

GET /api/errors
{
  "ok": true,
  "docs_base": "https://your-api.com/errors",
  "summary": {
    "total": 10,
    "retryable": 5,
    "non_retryable": 5,
    "by_action": {
      "fix_input_format": 2,
      "wait_and_retry": 1,
      "contact_operator": 1,
      "reauthenticate": 3,
      "no_action_possible": 2,
      "pick_different_value": 1
    }
  },
  "codes": [
    "ALREADY_EXISTS",
    "FORBIDDEN",
    "INTERNAL",
    "INVALID_FIELD",
    "INVALID_KEY",
    "MISSING_AUTH",
    "MISSING_FIELD",
    "NOT_FOUND",
    "RATE_LIMITED",
    "STATE_CONFLICT"
  ]
}

Or for a single code:

GET /api/errors?code=NOT_FOUND
{
  "ok": true,
  "code": "NOT_FOUND",
  "feedback": {
    "expected": "{resource} with id {got} not found",
    "current": {},
    "suggestion": null,
    "retryable": false,
    "action": "no_action_possible",
    "docs_url": "https://your-api.com/errors/NOT_FOUND"
  },
  "template_params": ["got", "resource"]
}

This lets the agent learn the entire error contract at the start of a session, rather than discovering errors one at a time.


Testing

The framework ships with a registry self-test:

npm test

The selftest asserts:

  1. Every action is from FEEDBACK_ACTIONS — catches typos in registry entries.
  2. Every retryable is a boolean — catches undefined/null mistakes.
  3. Every expected_template has a {param} placeholder — catches generic empty strings.
  4. Every suggestion_template is object or null — catches accidental strings.
  5. Every docs_url is an https URL — catches typos.
  6. Every suggestion-template key is referenced in the expected template — catches orphaned suggestions.
  7. Every code used in src/validators.js is in the registry — catches fail('NEW_CODE', ...) without a registry entry.
  8. Every code in the registry is used (or marked as unusedAllowed) — catches dead codes.

Run this in CI. Treat any failure as a blocker.


Adding new error codes

When you need a new code:

  1. Decide the action. Is this a retry_with_*? A fix_*? A pick_different_*? A *_and_retry? Look at the action vocabulary above. If none fit, your error is genuinely novel — propose a new action in a GitHub issue first (adding actions is a breaking change for agents).

  2. Decide if it's retryable. Will applying the suggestion likely succeed? If yes, retryable: true. If the suggestion is a placeholder for human review, retryable: false.

  3. Write the expected_template. Use {param} placeholders for any value the agent needs to know. Keep it human + LLM readable; the agent will read it directly when planning a retry.

  4. Write the suggestion_template (optional). If the server can compute a fix automatically, put it here. Use {param} for any computed value (the framework will render it from the params). null if no auto-fix is possible.

  5. Add the entry to FEEDBACK_REGISTRY in src/agentic-feedback.js. Match the existing format.

  6. Use the new code in your validators. Either:

    return failWithFeedback('MY_NEW_CODE', 'human message', { ...params });

    or:

    throw makeErrorWithFeedback('MY_NEW_CODE', 'human message', { ...params });
  7. Run npm test. The selftest will assert the code is in the registry, the action is valid, the templates are well-formed, and every used code has a matching registry entry.

  8. Document the code. Add a page at docs_url so the human reviewer can read the canonical explanation.


Language support

The wire format is just JSON, so the protocol works in any language. This repo ships with:

  • Node.js reference implementation (the framework itself, plus the Express adapter)
  • Node.js example client (example/agent-client.js)
  • Python example client (example/agent_client.py)

To port the protocol to another language:

  1. Implement the registry. Map error codes to their templates and actions. Keep the shape (action, retryable, expected_template, suggestion_template, docs_url) identical.

  2. Implement the template renderer. {param} → value lookup, recursive object rendering. ~30 lines of code in any language.

  3. Implement buildFeedback(code, params). Takes a code and dynamic params, returns the rendered envelope.

  4. Implement the agent-side dispatcher. A map of action values to handler functions. The set of actions is small; the handlers are the integration with your agent's planning layer.

That's it. The protocol is the wire format; the framework is just the reference implementation in Node.


FAQ

Q: Why not just use HTTP status codes?

Status codes are too coarse. A 400 can mean "missing field", "wrong type", "wrong format", "wrong value", "wrong network", "rate limited", "server bug", etc. The agent has to guess which one, and its guess is often wrong.

The framework doesn't replace HTTP status codes; it adds a structured body to them. A well-designed endpoint still returns 400 for a validation error and 500 for a server bug — but the body tells the agent which validation error, with a suggestion for the fix.

Q: Why not just return a more verbose error message?

Verbose messages help humans but don't help agents. The LLM still has to parse the message, guess the right action, and re-engineer the request. The framework provides structure that bypasses all of that.

A verbose message also tends to drift over time (someone "improves" the wording, and now the agent's regex-based parser breaks). A structured envelope is stable; the message is human-only.

Q: Doesn't this leak server internals?

expected and current are exactly the information a developer needs to debug the agent's call. They don't expose server internals (database structure, internal endpoints, etc.) — just the shape of the failed request.

If you have sensitive information (auth tokens, secrets in the request), make sure your validators don't echo them. The framework doesn't filter current automatically; you control what goes in.

Q: How do I version the framework?

The protocol is the wire format. Major version bumps happen when:

  • An action is added or removed
  • The envelope shape changes
  • The semantics of retryable change

Minor version bumps happen when:

  • New registry entries are added
  • New helpers are added (without breaking existing ones)

The wire format is JSON, so adding new fields is non-breaking (old clients ignore unknown fields). Removing or renaming a field is a major version bump.

Q: Why the ok: true/false convention?

Two reasons:

  1. Unambiguous discrimination. A parser can check ok and know whether to look for feedback or for the data fields. There's no "what if the response is empty" or "what if the body is an array" edge case.

  2. Symmetry. Every response has the same shape. A Promise<SuccessResponse | ErrorResponse> is exactly the type signature; the agent doesn't need to write special-case code for each endpoint.

Q: Can I use this with REST/HTTP/JSON-RPC/gRPC?

Yes, as long as the error response body is JSON. For gRPC, you'd adapt the envelope into a google.rpc.Status details field; for JSON-RPC, put it in the error.data field. The protocol doesn't care about the transport.

Q: What about WebSocket / streaming?

The framework assumes a single request → single response shape. For streaming, you'd extend the protocol: each event in the stream includes the same envelope structure on failure, plus an end_of_stream marker.

This is left as future work; the current framework is for the request/response case.

Q: Can I disable the envelope for non-agent clients?

Yes. Pass ?feedback=none in the query string and the server will return only {ok, error, code} without the envelope. This is for backward compatibility with legacy clients.

GET /api/things/99999?feedback=none
{
  "ok": false,
  "error": "thing 99999 not found",
  "code": "NOT_FOUND"
}

Q: How does this compare to RFC 7807 (Problem Details for HTTP APIs)?

RFC 7807 (and the follow-up RFC 9457) is a great standard for human-readable error responses. It standardizes fields like type, title, detail, instance.

The Agentic Feedback Framework is agent-readable. It standardizes fields like code, action, suggestion, retryable — things an LLM-driven client needs but a human developer doesn't.

You can use both side by side. RFC 7807 fields go in the top level; the framework envelope goes in the feedback key. Or, if you only care about agents, you can drop RFC 7807 and use just the framework.

Q: How does this compare to OpenAI's function-calling error format?

OpenAI's function-calling returns a free-text message field. The agent (which is the LLM itself) parses the message and decides what to do.

The Agentic Feedback Framework is for the next step: when the LLM-driven agent makes an HTTP call to your API, the response has the structured envelope. The LLM (or its client wrapper) doesn't have to parse free text; it dispatches on action directly.


License and contributing

This framework is MIT-licensed. See LICENSE.

Contributions are welcome. See CONTRIBUTING.md for the workflow for adding new error codes, new actions, or new language bindings.

This framework was originally developed for an agent-first application that handled thousands of agent requests per day with near-zero debugging cycles from the operator. It proved its value there and is now extracted here as a standalone, language-agnostic protocol. See CHANGELOG.md for the version history.

If you build something with this framework, open a PR and add it to the examples section. The point of the framework is to make agentic APIs safer and more debuggable, and the more reference implementations we have, the better.

Whether the agent is interacting with a decentralized exchange, a marketplace, a data service, or another autonomous system, failures become recoverable events instead of dead ends. The application provides the context, the correction, and the recovery path. That is the consistent interaction layer this framework exists to enable.

About

No description, website, or topics provided.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages