Build an enterprise AI agent: architecture and implementation guide

Applies to: Noviz AI platform, the open-source ERP AI Agent, ERPNext on Frappe Framework 15 and 16 · Audience: solution architects, ERP developers, technical decision makers · Reading time: about 60 minutes · Last updated 8 October 2026

This guide explains how to design and build an AI agent that works inside an enterprise system of record. It describes the architecture Noviz AI runs in production, the reasons behind each design decision, and the parts you can reuse from our open-source agent. Code samples are short, illustrative excerpts that show how each part is built. They contain no customer business rules and no production prompts.

Key takeaways

  • The language model reasons; the ERP system decides. Rules, permissions and workflows stay in the ERP.
  • The agent acts only through typed tools, and every tool call passes one gateway that checks permissions.
  • Every call runs as the signed-in user, so the ERP's own permissions and audit trail apply unchanged.
  • Token cost and accuracy are architecture problems: tool filtering, layered prompts and server-side analytics solve them.
  • A small engine you own is easier to make reliable than a general-purpose agent framework.

Overview

Most organisations already run their operations on a system of record: an ERP, a CRM, an HR suite. An enterprise agent is valuable when it can work with that system the way a trained employee does: read the right records, understand what they mean, prepare transactions, follow the approval process, and explain results. It is risky when it bypasses permissions, invents numbers, or changes data without consent.

The architecture in this guide is designed to keep the value and remove the risk. It has been built and refined on live ERPNext installations, extended to Microsoft Dynamics 365 Business Central, and is published in part as open source so that teams can study and reuse it.

PartWhat it covers
1. ConceptsWhat distinguishes an enterprise agent from a chatbot, and the principles every design choice follows.
2. ArchitectureThe reference architecture, the two deployment models, and the path of a single request.
3. ComponentsEach building block in detail: plugin, connector, tools, permissions, tool filtering, prompts, context, LLM connection, reasoning loop, rendering, analytics, RAG, scanning and safety.
4. DecisionsWhy we built our own engine instead of using LangChain or a similar framework, the libraries we chose, and how the system scales.
5. Build it yourselfHow to install the open-source agent, how the code is organised, and how to extend it.

Part 1. Concepts

What an enterprise agent is

A chatbot answers questions from text. An enterprise agent completes work in a business system. It reads live records, creates and updates documents, submits transactions, moves documents through approval workflows, builds analyses and reports, and processes incoming documents such as supplier invoices. It does this on behalf of a specific user, within that user's permissions.

CapabilityChatbotEnterprise agent
Source of answersModel knowledge and uploaded textLive records from the system of record, plus indexed process knowledge
ActionsNoneCreate, update, submit, approve, email, generate documents
IdentityAnonymous or a shared accountThe signed-in user; the ERP sees the real person
NumbersEstimated by the modelComputed exactly by the server; the model only describes them
RulesWritten into promptsEnforced by the ERP; the agent reads the outcome
AuditChat transcriptERP audit trail plus a structured interaction log

Design principles

Every component described in this guide follows five principles. When two designs compete, these decide.

  1. We are designing an AI, not an ERP. The ERP owns the business: data, validation rules, standard flows, approval workflows. The agent does not reimplement any of it.
  2. The model reasons from tool results. The agent decides its next step from what the ERP returned: the record, its state, the allowed transitions, or a refusal. There are no hidden conditions in the server or in prompts that decide on the model's behalf.
  3. The server provides tools and safety only. It offers good generic tools, enforces permissions and the customer's subscription tier, asks for confirmation before anything is committed, and passes results through unchanged.
  4. A refusal is information. When the ERP refuses an action (for example, insufficient stock or a closed accounting period), the agent shows the ERP's own message. It does not rephrase it into a guess.
  5. Exact numbers come from code. Totals, counts, averages and trends are computed by the server or the analytics engine. The model never adds up rows.

Important

Avoid encoding business rules in prompts ("never sell below cost", "orders above 10,000 need approval"). Prompt rules are not enforced, they drift from the real configuration, and they differ per customer. Configure them in the ERP, as validations, pricing rules or approval workflows, and let the agent read the result.

Part 2. Architecture

Reference architecture

The architecture has four layers: the channel where users work, the agent engine that reasons and orchestrates, the capability services the engine calls (language model, knowledge, analytics), and the system of record that holds and governs the data.

Reference architecture of an enterprise agent Users work in the ERP or a web app. Requests go to the agent engine, which selects tools, builds the prompt, calls the language model, and executes tool calls through a permission gateway and a connector to the ERP. The engine also uses a knowledge store and an analytics engine. Channel ERPNext Desk chat (plugin) · Business Central add-in · Web app · Email inbox Agent engine Auth and sessionuser, roles, tools Tool filteringmodules for this request Prompt and contextlayered prompt, memory Reasoning loopsteps, guards, limits Gatewaypermission, confirm Rendererstable, cards, chart, PDF Capability services Language model (any provider) Knowledge store (pgvector) Analytics engine (Python) Connector and system of record Canonical entities → native doctypes and fields · runs as the signed-in user ERPNext: data, validations, permissions, workflows, reports, print formats Agent data Sessions, conversations Interaction log, settings
Figure 1. Reference architecture. Arrows show the direction of requests; results flow back the same way.

The responsibilities are deliberately narrow:

  • Channel. Collects the user's request and displays structured results. It holds no business logic.
  • Agent engine. Authenticates, selects the tools relevant to the request, assembles the prompt and context, runs the reasoning loop, and executes tool calls through the gateway.
  • Gateway. The single point through which every tool call passes. It checks that the user may call the tool and that writes are confirmed.
  • Connector. Translates canonical entities (such as sales_order) into the ERP's native objects and fields, and calls the ERP as the signed-in user.
  • Capability services. The language model, the knowledge store for retrieval, and the analytics engine for exact calculations.
  • System of record. Decides everything that matters: whether data exists, whether an action is permitted, whether it is valid.

Deployment models

The same engine supports two deployment models. Choose by who operates the agent and where the data must stay.

Self-hosted (open source)Hosted relay with ERP plugin (Noviz AI)
Where the engine runsYour server, one deployment per organisationThe Noviz platform, shared by many organisations
Connection to the ERPThe engine calls the ERP's REST API directlyA plugin inside the ERP executes calls; the relay never connects to the ERP
ERP credentialsStored by the engine (encrypted per user)Never leave the ERP
User interfaceSeparate web app (agent and admin consoles)Inside the ERP (Desk page, Business Central page)
Tenancy, billing, tiersNot applicablePer-tenant keys, tiers, token budgets, usage metering
Best forTeams that want full control and can operate the stackOrganisations that want the agent working inside the ERP with no infrastructure

Request lifecycle

The following sequence describes one request, for example "Which customer invoices are overdue by more than 30 days?"

  1. Authenticate. The channel sends the request with the user's identity. The engine resolves the user's ERP roles and, from them, the list of tools that user may call.
  2. Load the conversation. Previous turns are loaded and trimmed to a fixed history budget.
  3. Select tools. A keyword router maps the request to business modules (here: accounting). Only the tools for those modules, plus a small always-available spine, are offered to the model.
  4. Assemble the prompt. A stable core, the sections for the matched modules, the rule blocks of the selected tools, and a short session block (date, company, user, roles).
  5. Add context. Recent conversation state and the closest passages from the knowledge store.
  6. Call the model. The model returns either an answer or one or more tool calls.
  7. Execute through the gateway. Each tool call is checked against the user's allowed tools, then executed by the connector as the user. Writes stop here for confirmation.
  8. Loop. Results are appended to the conversation and the model decides the next step. Guards stop loops that cannot converge.
  9. Render. The final response is a structured object (text, table, cards, chart specification, file link) that every channel renders the same way.
  10. Log. The prompt, tools called, outcome, latency and later user feedback are written to the interaction log.

Part 3. Components

ERP plugin and relay protocol

In the hosted model, the agent runs inside the ERP through a small plugin: a Frappe app for ERPNext, an AL extension for Business Central. The plugin provides the chat page and settings, and, most importantly, it executes the agent's data calls inside the ERP session of the user who asked. The relay (the hosted agent engine) never receives database credentials and never opens a connection to the customer's ERP.

A turn is an exchange of structured messages:

  1. The plugin sends the user's message, identity and roles to the relay, authenticated with the organisation's API key.
  2. The relay reasons and, instead of calling the ERP itself, returns a list of call specifications.
  3. The plugin executes each call with the ERP's own API, as the current user, and posts the results back.
  4. The relay continues the reasoning loop until it returns the final response.
// Relay to plugin: "run these calls for me, as this user"
{
  "turn_id": "6f1c0b2e-...",
  "status": "calls",
  "calls": [
    { "id": "c1", "method": "get_list", "doctype": "Sales Invoice",
      "fields": ["name", "customer", "due_date", "outstanding_amount"],
      "filters": [["docstatus", "=", 1], ["outstanding_amount", ">", 0],
                  ["due_date", "<", "2026-09-08"]],
      "order_by": "due_date asc", "limit": 50 }
  ]
}

// Plugin to relay: the results, exactly as ERPNext returned them
{ "turn_id": "6f1c0b2e-...", "results": [ { "id": "c1", "ok": true, "rows": [ ... ] } ] }

This design has three consequences that matter to customers:

  • Permissions are native. Because ERPNext executes the call in the user's session, its role permissions, user permissions and field restrictions apply automatically.
  • Audit is native. Documents created by the agent show the real user as owner in ERPNext's own history.
  • No inbound access. The customer opens no firewall ports and shares no credentials. The plugin makes outbound HTTPS calls only.

Note

Keep the plugin thin. It should execute calls and display results, nothing more. All reasoning, prompts and tool logic stay in the engine, so improvements reach every installation without a plugin update.

Connector layer

The engine never uses ERP-specific names. It works with canonical entities (customer, sales_order, leave_application) and canonical fields. A connector maps them to the ERP's native objects. This is what allows one engine to serve ERPNext and Business Central, and later other systems, without changes to tools, prompts or the gateway.

// Every connector implements one interface (excerpt from the open-source agent)
export interface SystemConnector {
  loginWithPassword(identifier: string, password: string): Promise<UserCredential>;
  getUserRoles(identifier: string): Promise<string[]>;
  list(entityKey: string, credential: UserCredential,
       params?: { filters?: Record<string, any>; limit?: number; sortBy?: string }): Promise<any[]>;
  get(entityKey: string, credential: UserCredential, id: string): Promise<any>;
  create(entityKey: string, credential: UserCredential, data: Record<string, any>): Promise<any>;
  update(entityKey: string, credential: UserCredential, id: string, data: Record<string, any>): Promise<any>;
  submit(entityKey: string, credential: UserCredential, id: string): Promise<any>;
  aggregate(entityKey: string, credential: UserCredential, params: AggregateParams): Promise<any>;
  runReport(reportKey: string, credential: UserCredential, filters?: Record<string, any>): Promise<any[]>;
}

// The ERPNext mapping: canonical entity -> native doctype and fields
sales_order: {
  doctype: "Sales Order",
  fieldMap: { id: "name", customer: "customer", status: "status",
              total: "grand_total", date: "transaction_date" },
},
sales_invoice: {
  doctype: "Sales Invoice",
  fieldMap: { id: "name", customer: "customer", status: "status", total: "grand_total",
              outstanding_amount: "outstanding_amount", due_date: "due_date", date: "posting_date" },
},

Source: backend/src/core/types.ts, backend/src/erpnext/entityMaps/selling.ts in the open-source agent

The connector registers itself once at start-up. Engine code asks the registry for the active connector and never imports ERP files directly. Adding a new ERP therefore means writing a connector and its entity maps, with no change to the engine.

Tools and module registry

The model can act only through tools. A tool has a unique name, a precise description written for the model, a JSON schema for its arguments, and a handler. Tools are grouped into modules, and modules are registered in a central registry at start-up.

export interface ToolDefinition {
  name: string;               // "<module>.<action>", globally unique
  description: string;        // written for the model: when to use it, what it returns
  module: string;
  parameters?: object;        // JSON schema for the arguments
  handler: (args: any, session: Session) => Promise<any>;
  promptRules?: string[];     // guidance injected only when this tool is offered
}

export const knowledgeModule: MCPModule = {
  name: "knowledge",
  description: "ERP process knowledge and live approval workflow",
  tools: [ knowledgeSearchTool, workflowDescribeTool ],
};

Source: backend/src/core/types.ts, backend/src/modules/knowledge/index.ts in the open-source agent

Most tools are not written by hand. Three factories generate them from declarative configuration:

FactoryInputGenerated tools
Entity factoryAn entity configuration: canonical fields, which operations are exposed<entity>.list, .get, .create, .update, .submit
Report factoryA report configuration: canonical filters mapped to the ERP reportreport.<name> tools that run the ERP's own report
Workflow factoryA state transition definitionOne gated tool per transition

Hand-written modules cover what configuration cannot: analytics, charts, report and PDF generation, schema discovery, inbox actions, document scanning and knowledge retrieval.

Tip

Treat tool descriptions as part of the product. A precise description ("returns at most 50 rows; use analytics.aggregate for totals") prevents more errors than any prompt instruction, because the model reads it exactly when it chooses the tool.

Identity and permissions

Permissions are enforced in two independent layers. Both must allow an action.

  1. Tool permission (agent layer). At sign-in, the user's ERP roles are resolved into the list of tools that user may call. A sales user is never offered payroll tools. The model cannot call a tool it was not offered, and the gateway rejects it if it tries.
  2. Data permission (ERP layer). Every call runs as the user, so the ERP applies its own role permissions, user permissions and document-level restrictions. If the agent layer is misconfigured, the ERP still refuses.
// The single enforcement and execution point for every tool call
export async function callTool(session: Session, toolName: string, args: any) {
  const allowed = session.allowed_tools.includes("*") || session.allowed_tools.includes(toolName);
  if (!allowed) throw new ToolNotAllowedError(toolName);

  const tool = moduleRegistry.findTool(toolName);
  if (!tool) throw new Error(`Unknown tool: ${toolName}`);

  return tool.handler(args, session);   // the handler calls the connector as session.credential
}

Source: backend/src/core/gateway.ts (business-rule hooks omitted) in the open-source agent

Nothing else in the engine may call a tool handler directly. A single entry point makes the permission model auditable: to verify it, you read one function.

Write confirmation. Any tool that creates, updates, submits, cancels or sends returns a proposal first: the exact document and values it will write. The user confirms in the interface, and only then does the engine execute the call. A tenant can also be limited to read-only access, in which case write tools are never offered.

Tool filtering

An ERP agent can have several hundred tools: five operations for each of about 80 entities, plus reports and utilities. Offering all of them on every request is slow, expensive, and less accurate, because the model has more ways to choose wrongly. Model providers also impose hard limits on the number of tools per request.

We filter tools per request in three stages:

  1. Role filter. Only tools the user is permitted to call.
  2. Module router. A deterministic keyword match maps the request to business modules (selling, buying, stock, accounting, HR and so on). Related modules are widened together: a CRM question that mentions invoices also loads selling. Only the read tools of matched modules are offered by default.
  3. Spine and discovery. A small spine of tools is always available, including tools.search, which lets the model find and enable any other permitted tool by describing what it needs. Tools found this way stay available for the rest of the turn.
const RELAY_SPINE_TOOLS = new Set([
  "tools.search",
  "data_table.list",
  "data_table.search_schema",
  "database_engine.execute_query",
  "report.find",
  "analytics.catalog",
]);

export function selectRelayTools(tools: ToolDefinition[], prompt: string): ToolDefinition[] {
  const modules = relayModulesFor(prompt);
  const wantsCharts = detectExplicitModules(prompt).has("analytics");
  return tools.filter((t) => {
    if (RELAY_SPINE_TOOLS.has(t.name) || t.module === "context") return true;
    if (wantsCharts && (t.name === "analytics.aggregate" || t.name === "chart.build")) return true;
    if (modules.size > 0 && (t.name.endsWith(".list") || t.name.endsWith(".get")) && toolMatchesBusinessModule(t, modules)) {
      return true;
    }
    return false;
  });
}

Source: backend/src/core/toolRelevanceFilter.ts in the open-source agent

The router is intentionally simple and deterministic. We evaluated model-based tool selection (an extra model call that picks tools) and keep it only for very broad roles such as system administrators. For department roles, the keyword router is faster, free, and predictable.

System prompt design

A single large system prompt is the most common cause of cost and inconsistency in agent projects. Every rule is paid for on every request, and rules for one domain interfere with another. We assemble the system prompt from layers instead:

LayerContentWhen included
CoreIdentity, operating principles, how to choose tools, display conventionsAlways; identical bytes on every request
Write layerHow to propose changes and wait for confirmationOnly when the user has write access
Module sectionsDomain guidance (for example, how selling documents relate)Only for modules the router matched
Tool rule blocksGuidance a specific tool needs (query discipline, analytics, reports)Injected when that tool is offered, at most once per conversation
Organisation policyFree-text policies the customer configuresNew conversations, for that organisation only
Session blockDate, company, user, rolesAlways; always last
export function buildSystemPrompt(prompt: string, canWrite: boolean, identity: AppIdentity, frappeUser: string, frappeRoles: string[]): string {
  const realModules = [...relayModulesFor(prompt)].filter((m) => BUSINESS_MODULE_KEYS.includes(m));

  const sections: string[] = [];

  if (realModules.length === 0) {
    sections.push(CORE_SYSTEM_PROMPT);
  } else {
    sections.push(THIN_CORE, AUTOFLOW_RULES); // AUTOFLOW: any real business-data turn wants the document-chain facts
  }

  if (canWrite) sections.push(WRITE_OPERATIONS);

  const seen = new Set<string>();
  for (const m of realModules) {
    for (const section of MODULE_PROMPT_SECTIONS[m] || []) {
      if (!seen.has(section)) {
        seen.add(section);
        sections.push(section);
      }
    }
  }

  sections.push(buildUserContext(identity, frappeUser, frappeRoles));
  return sections.filter(Boolean).join("\n\n");
}

Source: backend/src/systemPrompt/index.ts in the open-source agent

The fixed order is not cosmetic. Model providers cache the beginning of a prompt that is byte-for-byte identical across requests and charge less for it. Keeping the stable core first and the per-user session block last lets most of every request hit that cache.

Important

Treat prompts as production code. Change them through review, with the full test suite, and record why each rule exists. A rule added to fix one conversation can break ten others; a regression suite of real requests is the only reliable protection.

Where to add guidance

The open-source agent ships every prompt block empty. Each organisation writes its own guidance. The table shows where each kind of guidance goes; the examples are placeholders that show the shape, not recommended wording.

GuidanceWhere to add itExample
Who the agent issystemPrompt/core/identity.ts"You are the operations agent for Acme."
How to choose toolssystemPrompt/core/toolDiscovery.ts"For totals, use analytics.aggregate."
How to present resultssystemPrompt/core/display.ts"Show lists as tables."
Before any changesystemPrompt/core/writeOperations.ts"Show the values and wait for confirmation."
One business areasystemPrompt/modules/selling.ts (one file per module)"A quotation becomes a sales order."
When a module is loadedMODULE_KEYWORDS in systemPrompt/modules/index.tsselling: ["quotation", "sales order"]
Guidance for one toolThe tool’s promptRulespromptRules: [REPORT_RULES]
Report toolssystemPrompt/core/reports.ts"Use reports for files."
Per-message hintssystemPrompt/core/hints.ts and the *_HINT constants in core/"Use a relative date filter."
Your company’s policiesAdmin console, Policy documents (no code)A policy document, uploaded
How processes workbackend/knowledge/ (no code)A Markdown file, indexed

A module section is one exported string, and the keywords decide when it is sent:

// systemPrompt/modules/selling.ts
export const SELLING_MODULE = "A quotation becomes a sales order.";

// systemPrompt/modules/index.ts
selling: ["quotation", "sales order"],

Placeholder wording. A request that mentions “quotation” loads the selling section and the selling tools together.

A tool rule travels with its tool, and is added to the prompt only when that tool is offered:

export const REPORT_RULES = "Use reports for files, not chat tables.";

{ name: "report.generate", /* ... */ promptRules: [REPORT_RULES] }

Placeholder wording. The rule is sent once per conversation, the first time the tool appears.

Important

Keep business rules out of every block. Limits, approvals and validations belong in ERPNext; a prompt only explains how to use the tools.

Context management

Context is everything the model sees besides the request and the tools. We organise it in three tiers by cost:

TierSourceHow it is used
HotSession cache: the current conversation, recently referenced records, the last result setAlways included, within a character budget
WarmVector store: knowledge documents, organisation policiesThe closest passages are added automatically
ColdAnything else in the ERP or the knowledge storeFetched only when the model calls a tool for it

Three practices keep context accurate and affordable:

  • Budgets per provider. Each context source has a character budget, so one large source cannot crowd out the others.
  • Bounded history. Conversations keep a fixed number of previous user turns. Older turns are dropped, not summarised by the model, which avoids silent loss of detail.
  • Large results stay on the server. When a tool returns many rows, the model receives a page and a summary. The full set is kept server-side for paging, export and analysis.

LLM connection

The engine talks to the language model through one provider interface. Any service that implements the widely used chat-completions format with tool calling can be used: OpenAI, DeepSeek, Azure OpenAI, or a self-hosted model behind a compatible server.

export interface LLMProvider {
  chat(request: {
    messages: ChatMessage[];
    tools?: ToolSchema[];
    temperature?: number;
    maxTokens?: number;
  }): Promise<{ content: string | null; toolCalls: ToolCall[]; usage: TokenUsage }>;
}

Illustrative. The open-source implementation is backend/src/providers/llm/openaiProvider.ts

Operational practices that proved necessary in production:

  • Switchable at runtime. The base URL, model name and key are settings, not code. Switching provider takes seconds and needs no deployment.
  • Separate models per task. Reasoning, document vision and embeddings can use different providers and keys, so one provider's outage or limit does not stop everything.
  • Timeouts and retries. Every call has a timeout; transient errors are retried with back-off; a clear message is returned when the provider is unavailable.
  • Rate and budget gates. A shared token-rate gate keeps the service under the provider's limits, and a monthly token budget per organisation prevents runaway cost.
  • Usage recorded per request. Prompt, cached and completion tokens are logged for every turn, which is how cost problems are found.

Reasoning loop and workers

The reasoning loop is a plain loop: select the tools for this step, inject their rule blocks, call the model, check the proposed calls, execute them, append the results, repeat. Tools the model finds with tools.search stay available for the rest of the turn. We keep it as ordinary functions rather than a graph or framework, because each step must be easy to read, test and trace.

const toolsForThisStep = () => {
  const selected = selectRelayTools(allowedTools, relevanceSearchText);
  const selectedNames = new Set(selected.map((t) => t.name));
  const forced = allowedTools.filter((t) => forcedToolNames.has(t.name) && !selectedNames.has(t.name));
  return [...forced, ...selected].slice(0, MAX_TOOLS_PER_REQUEST);
};

for (let i = 0; i < appConfig.llm.maxToolIterations; i++) {
  const tools = toolsForThisStep();
  for (const t of tools) for (const block of t.promptRules ?? []) injectPromptBlock(messages, block);
  const response = await this.llm.chat(messages, tools);
  if (response.tool_calls.length === 0) { finalText = response.content || ""; break; }
  messages.push({ role: "assistant", content: response.content || "", tool_calls: response.tool_calls });
  for (const call of response.tool_calls) {
    // guards: repeated calls, denied tools, step and time budgets ...
    const result = await callTool(session, call.name, call.arguments);
    if (call.name === "tools.search") {
      for (const name of foundToolNames(result).slice(0, TOOLS_SEARCH_FORCE_CAP)) forcedToolNames.add(name);
    }
    messages.push({ role: "tool", content: JSON.stringify(result), tool_call_id: call.id, name: call.name });
  }
}

Source: backend/src/core/reasoningEngine.ts (shortened) in the open-source agent

Guards. Models occasionally repeat a call that cannot succeed, such as paging through thousands of rows to produce a total. Guards detect these patterns and redirect the model to the right tool (an aggregate, or a PDF report). Each guard exists because of a specific, reproduced failure, and each has a test. Remove a guard only when the failure it covers is proven impossible.

Workers. The engine is stateless between steps: conversation and turn state live in the database and Redis. Several worker processes therefore serve requests in parallel behind one endpoint, and any worker can continue any turn. Long-running turns can move to a background queue so that the web request returns immediately and the channel polls for the result.

Rendering

The model does not produce HTML, and it does not draw charts. It chooses a response type, and the engine produces a structured object that every channel renders with the same components.

Response typeUsed forProduced by
TextExplanations, confirmations, refusals passed through from the ERPThe model
TableLists of records, with paging and links to the ERP documentTable renderer from tool results
CardsA single record, an email or notification with action buttonsCards renderer
Chart specificationBar, line, pie and combined dashboardsChart builder from exact aggregates
FilePDF reports, ERP print formats, exportsReport generator or the ERP's own print engine
ProposalA write awaiting confirmationThe gateway
{
  "type": "table",
  "title": "Overdue customer invoices",
  "columns": ["Invoice", "Customer", "Due date", "Outstanding"],
  "rows": [["ACC-SINV-2026-00412", "Gulf Traders LLC", "2026-08-14", 18450.00]],
  "links": { "Invoice": "/app/sales-invoice/{value}" },
  "page": { "offset": 0, "size": 20, "total": 37 }
}

Because the response is data, one design serves every surface: the ERPNext Desk page, a web app and the Business Central add-in show the same table, the same chart and the same confirmation dialog. Values in tables and charts come from tool results, so the model cannot alter a number on its way to the screen.

Analytics engine

Language models are unreliable at arithmetic over many rows and expensive when given thousands of records. Analytics therefore follows one rule: rows never go to the model. The engine computes the result and the model receives a compact summary to explain.

There are two levels:

  • Exact aggregates in the engine. Sums, counts, averages, minimum and maximum, percentages and group-by are computed by the connector, using the ERP's own count and aggregation where available and chunked fetching otherwise.
  • Named analyses in Python. Trends, growth, ageing, concentration, simple forecasting and anomaly detection run in a Python analytics engine with pandas, NumPy and SciPy. Each analysis returns KPIs, a chart specification and written observations.

The following example is the complete sample analysis shipped with the open-source agent: a monthly trend with growth and a one-month projection.

"""Monthly trend: the sample analysis of the Python analytics engine.

Input  (stdin, JSON): {"rows": [{"date": "2026-01-14", "total": 1200.0}, ...]}
Output (stdout, JSON): KPIs, a chart specification and observations.

The rows use canonical field names (date, total), the data contract every
connector maps its native fields to, so the analysis works for any ERP.
The rows never reach the language model; only this output does.
"""
import json
import sys

import numpy as np
import pandas as pd


def analyse(rows):
    df = pd.DataFrame(rows)
    if df.empty or "date" not in df or "total" not in df:
        return {"kpis": [], "chart": None, "observations": ["No records in the period."]}

    df["date"] = pd.to_datetime(df["date"])
    df["total"] = pd.to_numeric(df["total"], errors="coerce").fillna(0.0)
    monthly = df.set_index("date")["total"].resample("MS").sum()
    growth = monthly.pct_change().mul(100).round(1)

    # Linear projection for the next month (least squares over the series).
    if len(monthly) > 1:
        slope, intercept = np.polyfit(np.arange(len(monthly)), monthly.values, 1)
    else:
        slope, intercept = 0.0, float(monthly.iloc[0])
    projection = max(0.0, float(slope * len(monthly) + intercept))

    best = monthly.idxmax()
    observations = [
        f"Highest month: {best:%B %Y} at {monthly.max():,.2f}.",
        f"The trend is {'rising' if slope > 0 else 'falling' if slope < 0 else 'flat'} by about {abs(slope):,.2f} per month.",
    ]
    if growth.notna().any():
        observations.append(f"Latest month-on-month change: {growth.iloc[-1]:+.1f}%.")

    return {
        "kpis": [
            {"label": "Total", "value": round(float(monthly.sum()), 2)},
            {"label": "Average per month", "value": round(float(monthly.mean()), 2)},
            {"label": "Projected next month", "value": round(projection, 2)},
        ],
        "chart": {
            "type": "line",
            "labels": [d.strftime("%b %Y") for d in monthly.index],
            "series": [{"name": "Total", "values": [round(float(v), 2) for v in monthly.values]}],
        },
        "observations": observations,
    }


if __name__ == "__main__":
    print(json.dumps(analyse(json.load(sys.stdin).get("rows", []))))

Source: backend/analytics/monthly_trend.py in the open-source agent

The tool fetches the rows through the connector, as the user, keeps only the fields in the data contract, and hands them to the analysis:

handler: async (args, session) => {
  if (!TREND_ENTITIES.includes(args.entity)) throw new Error(`entity must be one of: ${TREND_ENTITIES.join(", ")}`);
  const rows = await systemConnector.list(args.entity, session.credential, {
    filters: { date: ["between", [args.from_date, args.to_date]], status: ["not in", ["Draft", "Cancelled"]] },
    limit: ROW_CAP,
    sortBy: "date",
    sortDir: "asc",
  });
  const result = await runPythonAnalysis("monthly_trend", rows.map((r) => ({ date: r.date, total: r.total })));
  return {
    entity: args.entity,
    period: { from: args.from_date, to: args.to_date },
    records_analysed: rows.length,
    ...(rows.length >= ROW_CAP ? { note: `Only the first ${ROW_CAP} records were analysed; narrow the date range for a complete result.` } : {}),
    ...result,
  };
},

Source: backend/src/modules/trends/index.ts in the open-source agent

The runner starts Python as a child process. In a single deployment this is sufficient; at scale, the same scripts run unchanged as a separate Python service.

import { spawn } from "child_process";
import path from "path";

/**
 * Runs one Python analysis (backend/analytics/<name>.py) as a child
 * process: rows in on stdin, JSON result out on stdout. Keeps the agent a
 * single deployment; a hosted platform can run the same scripts as a
 * separate Python service without changing them. The rows go to Python,
 * never to the language model; the model only sees the returned summary.
 */
export const ANALYTICS_DIR = path.resolve(__dirname, "../../analytics");

export function runPythonAnalysis(
  name: string,
  rows: object[],
  opts: { command?: string; args?: string[]; timeoutMs?: number } = {}
): Promise<any> {
  if (!/^[a-z0-9_]+$/.test(name)) return Promise.reject(new Error(`Invalid analysis name "${name}"`));
  const command = opts.command ?? (process.env.PYTHON_BIN || "python3");
  const args = opts.args ?? [path.join(ANALYTICS_DIR, `${name}.py`)];

  return new Promise((resolve, reject) => {
    const child = spawn(command, args, { timeout: opts.timeoutMs ?? 30000 });
    let out = "";
    let err = "";
    child.stdout.on("data", (d) => (out += d));
    child.stderr.on("data", (d) => (err += d));
    child.on("error", (e) => reject(new Error(`Could not start the analytics engine (${command}): ${e.message}`)));
    child.on("close", (code) => {
      if (code !== 0) return reject(new Error(`Analysis "${name}" failed: ${err.trim() || `exit code ${code}`}`));
      try {
        resolve(JSON.parse(out));
      } catch {
        reject(new Error(`Analysis "${name}" returned invalid JSON`));
      }
    });
    child.stdin.end(JSON.stringify({ rows }));
  });
}

Source: backend/src/core/pythonAnalysis.ts in the open-source agent

Note

Define a data contract for analyses: the column names each analysis expects. Each ERP connector maps its data to that contract, so the same Python analyses serve ERPNext, Business Central and future systems without modification.

Knowledge and RAG

Tools give the agent access to data. They do not tell it what the data means: that a Delivery Note reduces stock, that a submitted invoice cannot be edited, that a Purchase Order in "Pending Approval" is waiting for a specific role. Retrieval-augmented generation (RAG) supplies that understanding from documents, at the moment it is needed, without putting it into every prompt.

Knowledge sources

SourceExample contentHow it is kept current
ERP process documentationOrder to Cash, Procure to Pay, stock movements, period close, leave and expense flows, the document lifecycle (draft, submitted, cancelled, amended)Versioned files, re-indexed when changed
Organisation documentsStandard operating procedures, policies, contracts uploaded by an administratorRe-indexed on every edit; can be deactivated
Live workflow definitionsThe organisation's own approval workflows: states, transitions, which role may actRead from the ERP at the time of the request, never copied

Ingestion pipeline

  1. Parse. Each document (a Markdown file under backend/knowledge/) carries metadata: title, module (selling, buying, stock), and the document types it describes.
  2. Chunk by meaning. Documents are split at section headings, so a chunk covers one idea. Long sections are split further with a small overlap, so a sentence on a boundary keeps its context.
  3. Label. Each chunk is prefixed with "document title > section", so it remains understandable once separated from its file.
  4. Embed and store. Chunks are embedded and stored in PostgreSQL with the pgvector extension, linked to their source document. Re-indexing a document replaces only its own chunks.
  5. Skip unchanged files. A checksum per document avoids paying for embeddings that have not changed.
/** Splits a document into labelled retrieval chunks: one per "## " section,
 *  long sections further split by the shared policy chunker. */
export function chunkKnowledge(doc: KnowledgeFile): KnowledgeChunk[] {
  const sections: { name: string; text: string }[] = [];
  let current = { name: "Overview", text: "" };
  for (const line of doc.body.split("\n")) {
    const h2 = line.match(/^##\s+(.+)$/);
    if (h2) {
      if (current.text.trim()) sections.push(current);
      current = { name: h2[1].trim(), text: "" };
    } else if (!/^#\s+/.test(line)) {
      current.text += line + "\n";
    }
  }
  if (current.text.trim()) sections.push(current);

  const chunks: KnowledgeChunk[] = [];
  for (const s of sections) {
    const parts = chunkText(s.text.trim());
    parts.forEach((part, i) => {
      const label = `knowledge:${doc.slug}#${s.name}${parts.length > 1 ? ` (${i + 1}/${parts.length})` : ""}`;
      chunks.push({ label, content: `${doc.title} > ${s.name}\n\n${part}` });
    });
  }
  return chunks;
}

Source: backend/src/core/knowledgeStore.ts in the open-source agent

Retrieval

Knowledge reaches the model in two ways:

  • Automatically. The closest passages to the request are added to the warm context tier, within its budget.
  • On demand. The model calls knowledge.search when it needs to understand a process before acting, optionally filtered by module, and cites the source title in its answer.
select d.title, e.label, e.content, 1 - (e.embedding <=> $1) as score
from   context_embeddings e
join   knowledge_documents d on d.id = e.knowledge_document_id
where  d.active and ($2::text is null or d.module = $2 or d.module is null)
order  by e.embedding <=> $1
limit  $3;

Source: backend/src/core/knowledgeStore.ts (search) in the open-source agent

Workflow awareness

Static documents describe the standard process. The organisation's real approval process lives in the ERP and changes over time. A second tool, workflow.describe, reads it live: the configured states and transitions for a document type and, for a specific document, its current state and the actions available to the current user right now, as computed by the ERP itself.

{
  "doctype": "Purchase Order",
  "has_workflow": true,
  "states": [ { "state": "Pending Approval", "doc_status": "Draft" },
              { "state": "Approved", "doc_status": "Submitted" } ],
  "transitions": [ { "from": "Pending Approval", "action": "Approve", "to": "Approved",
                     "allowed_role": "Purchase Manager", "condition": null, "self_approval": false } ],
  "document": { "name": "PUR-ORD-2026-00087", "current_state": "Pending Approval",
                "docstatus": 0, "actions_for_you": [] }
}

Output of workflow.describe (backend/src/erpnext/workflow.ts). The workflow definition is metadata; the document and the actions available on it are read as the user, through ERPNext’s own get_transitions.

With both tools, the agent can answer "why can't I submit this order?" correctly: the knowledge base explains approval workflows in general, and the live workflow shows that this order needs a Purchase Manager, which the current user is not. No rule in the agent made that decision; the ERP did.

Important

Do not use RAG to store business data or business rules. Data must be read live through tools, with the user's permissions, or answers will be stale and may expose records the user cannot see. Rules must be enforced by the ERP. RAG is for meaning: how processes work and why.

Document scanning

Incoming documents (supplier invoices, purchase orders from customers, receipts) are a natural task for an agent. The pipeline is:

  1. The user uploads a file or takes a photo in the channel.
  2. A vision-capable model extracts the text and the structure: party, dates, line items, totals.
  3. The agent identifies the matching transaction type and looks up the party and items in the ERP.
  4. It prepares a draft document and presents it as a proposal. Nothing is written until the user confirms.

Safety and governance

ControlImplementation
Least privilegeTools per role; data permissions by the ERP; read-only tier available
Consent before commitEvery write is a proposal until the user confirms
Truthful errorsERP refusals are shown as returned
No inbound accessIn the hosted model, the plugin calls out; the ERP is never exposed
Credential protectionPer-user credentials encrypted at rest (self-hosted); never transmitted (hosted)
Abuse and cost limitsRate limits per organisation, token-rate gate, monthly token budgets
Data protectionCustomer mailboxes are read without modification; result sets stay server-side
AccountabilityERP audit trail as the real user; interaction log for every request

Observability and testing

An agent is only as reliable as the evidence about how it behaves. We rely on three instruments:

  • Interaction log. Every request records the prompt, the modules matched, the tools offered and called, the response type, latency, token usage and later user feedback. Improvements start from real traffic, not assumptions.
  • Automated tests. Our engine has more than 1,200 automated tests covering tool generation, permission resolution, prompt assembly, guards, renderers and connector mapping. Every change runs the full suite.
  • Regression discipline. Before changing a guard, a prompt layer or session handling, read the history of why it exists and prove the change cannot reopen a fixed failure.

Part 4. Decisions

Why we did not use LangChain

LangChain, LangGraph, CrewAI, AutoGen and Semantic Kernel are useful for prototypes and for teams that need many integrations quickly. We evaluated them and chose to build a small engine of our own. The reasons are specific to enterprise agents that operate on a system of record.

ConcernWith a general frameworkWith our own engine
Prompt controlPrompts are assembled by framework templates; the exact bytes sent are hard to see and change between versionsWe control every byte and its order, which enables prompt caching and reproducible behaviour
Token costAgents typically receive all tools and long default instructionsPer-request tool filtering and layered prompts reduce tokens substantially
PermissionsTool access is an application concern added around the frameworkOne gateway is the only path to a tool; permissions are part of the core design
Deterministic guardsHard to insert between "model proposes" and "tool executes" without fighting abstractionsGuards are plain, named, tested functions in the loop
DebuggingFailures surface deep inside framework call stacksA short loop; each step is visible in the interaction log
Dependency stabilityFrequent breaking changes; large transitive dependency treeA small set of mature libraries that change rarely
Multi-agent patternsSeveral agents talking to each other multiply token use and latencyOne agent with many tools; specialist agents only where a task truly needs them
ScaleFrameworks organise code; they do not add capacityCapacity comes from stateless workers, caching and provider limits, which we manage directly

The real limits of an agent in production are provider rate limits, cost per request, accuracy, and database load. None of these is solved by a framework. They are solved by architecture: fewer tokens per request, exact computation outside the model, caching, and clear control over every step.

When a framework is the right choice

For a proof of concept, a research prototype, or an application that mainly connects many third-party data sources, a framework can save weeks. If you start with one, keep your permission gateway, connector and tool definitions in your own code, so you can replace the orchestration later without rewriting the system.

Libraries and services

AreaChoiceWhy
Engine runtimeNode.js with TypeScript, ExpressStrong typing for tool schemas and connectors; excellent I/O concurrency for API-bound work
DatabasePostgreSQL with pgvector (pg)One database for sessions, logs, settings and vectors; no separate vector service to operate
Cache and coordinationRedis (ioredis)Response cache, rate limits and the shared token-rate gate across workers
ERP and model HTTPaxiosTimeouts, interceptors and clear error handling for ERP and model calls
AuthenticationjsonwebtokenSigned session tokens for the channels
Documentspdfkit, pdf-parse, mammoth, multerGenerate PDF reports; extract text from uploaded PDF and Word documents
EmailnodemailerOutbound mail for reports and notifications
AnalyticsPython with pandas, NumPy, SciPyThe standard toolset for analysis, forecasting and anomaly detection
InterfacesReact with Vite; Frappe Desk page; AL control add-inOne response format rendered natively on each surface
TestsJest with ts-jestFast unit tests for every component of the engine

Scaling

  • Modular monolith first. The engine is one service with clear modules (core, connector, tools, prompts, renderers). Split a part into its own service only when it has a different resource profile, language or failure mode. Analytics (Python) and knowledge (vector ingestion) qualify; the reasoning hot path does not.
  • Stateless workers. Run several worker processes; turn state lives in PostgreSQL and Redis.
  • Response caching. Identical first questions are served from an exact cache; very similar questions from a semantic cache. Data-dependent answers are cached briefly and scoped per organisation.
  • Background turns. Long analyses and reports run in a queue so web requests stay fast.
  • Provider capacity. Monitor token usage per minute; raise provider tiers or distribute across keys before limits are reached.

Part 5. Build it yourself

Get the open-source agent

The ERP AI Agent is our open-source, self-hosted agent for ERPNext. It contains the engine described in this guide: the connector, the module registry and factories, the gateway, role-based tool permissions, context tiers, the analytics toolbox, renderers, report and PDF generation, document scanning, the knowledge base with live workflow awareness, a Python analytics sample, and the agent and admin web apps. It is licensed under AGPL-3.0.

github.com/Zeetech-soft-solution/erp-ai-agent

Prerequisites

  • An ERPNext instance (version 15 or 16) and an API key for a service user
  • Node.js 20 or later
  • PostgreSQL 15 or later with the vector extension
  • An API key for an OpenAI-compatible model provider

Install

git clone https://github.com/Zeetech-soft-solution/erp-ai-agent.git
cd erp-ai-agent/backend
npm install
cp .env.example .env          # ERPNext URL and key, LLM key, DATABASE_URL, CREDENTIAL_ENCRYPTION_KEY

# Create the schema: run every migration in filename order
for f in src/db/migrations/*.sql; do psql "$DATABASE_URL" -f "$f"; done

npm test                      # the full test suite

# Optional: knowledge base and Python analytics
npm run knowledge:index                          # embed backend/knowledge/**/*.md
pip install -r analytics/requirements.txt        # pandas and numpy for analytics/*.py

npm run dev                   # API on the port set in .env

cd ../frontend/agent && npm install && cp .env.example .env && npm run dev
cd ../admin && npm install && cp .env.example .env && npm run dev

Sign in to the agent app with an ERPNext user. The tools you see are those your ERPNext roles allow. The repository's docs/INSTALL.md describes each step in detail, including installing ERPNext itself, and docs/SAMPLE_PROMPTS.md lists requests to try.

Project structure

backend/src/
  core/            engine: gateway, reasoning loop, module registry, factories,
                   context assembly, tool filtering, renderers registry, stores
  config/          canonical entities, role policy, reports, workflows, settings
  erpnext/         the ERPNext connector: client, entity maps, report map
  modules/         hand-written tools: analytics, trends, chart, reports, documents,
                   schema, tool discovery, inbox actions, context, knowledge
  providers/       LLM, embeddings, vision, context tiers
  renderers/       table, cards, chart
  systemPrompt/    layered prompt: core blocks and module sections
  routes/          auth, agent, admin, tools, policy documents, webhooks
  db/migrations/   PostgreSQL schema
backend/knowledge/ process documents for the knowledge base (RAG)
backend/analytics/ Python analyses, run as a child process
frontend/
  agent/           the user-facing agent app (React)
  admin/           the administration console (React)

Extend the agent

TaskWhat to change
Add an ERP entityAdd an entity configuration (canonical fields and operations) and its mapping in the connector's entity map. The factory generates the tools.
Grant it to a roleAdd the generated tool names to that role in the role policy.
Add a custom toolWrite a module with a tool definition and handler; register it in the bootstrap list. The gateway applies permissions automatically.
Add process knowledgeAdd a Markdown file with title, module and sections under backend/knowledge/ and run npm run knowledge:index.
Add an analysisAdd a script under backend/analytics/ that reads rows in the data contract and prints KPIs, a chart and observations; expose it with a tool such as analytics.monthly_trend.
Connect another ERPImplement the connector interface and entity maps for that system. Tools, prompts, gateway and interfaces stay unchanged.
Use another model providerSet the base URL, model and key; any chat-completions provider with tool calling works.

Production checklist

  • Every tool call passes one gateway; no handler is called directly.
  • Every ERP call runs as the signed-in user; a service account is used only for role lookup and metadata.
  • Writes require explicit confirmation; read-only access is available.
  • No business rules in prompts; ERP refusals are shown as returned.
  • Totals and analyses are computed in code; rows are not sent to the model.
  • Tools are filtered per request; prompts are layered with a stable prefix.
  • Knowledge documents explain processes; live data and workflows are read through tools.
  • Timeouts, retries, rate limits and token budgets are configured.
  • Every request is logged with tools, tokens, latency and feedback.
  • The full test suite runs on every change; guards have tests that name the failure they prevent.

Next steps

Questions or corrections: support@noviz.in