How to build an enterprise AI agent

Applies to: enterprise systems of record (ERP, CRM, HR, finance, service management, supply chain) · Reference implementation: the open-source Noviz agent on ERPNext · Audience: enterprise architects, solution architects, platform engineers, technical decision makers · Reading time: about 90 minutes · Last updated 9 October 2026

This guide describes how to design, build and operate an AI agent that works inside the systems an enterprise already runs. It is written from the experience of building the Noviz platform, which serves agents for ERPNext and Microsoft Dynamics 365 Business Central from one common engine, and it explains each architectural decision together with the alternatives we considered and what production use taught us. The examples come from our open-source reference implementation, which uses ERPNext as its system of record; the architecture itself applies to any business system with an API.

In this guide

  • A reference architecture for enterprise agents, with deployment topologies and the path of a single request.
  • Sixteen building blocks, each with its purpose, the design we chose, the trade-offs and working code.
  • How to run an agent in production: observability, security, cost, scaling and change management.
  • The decisions that shape the platform, including why we built our own engine rather than using an agent framework.
  • How to start from the open-source agent and extend it to your own systems.

Overview

Enterprises run on systems of record: the ERP that holds orders, stock and ledgers; the CRM that holds customers and pipelines; HR, payroll, procurement, service desks and document stores. Each of these systems already contains the organisation’s data, its business rules, its permissions and its approval processes. An enterprise agent is valuable when it can work with those systems the way an experienced employee does: find the right information, understand what it means, prepare and complete transactions, follow the approval process and explain the outcome.

The difficulty is not making a language model talk about business data. It is making an agent that is correct, that never exceeds the permissions of the person using it, that keeps cost predictable at scale, and that remains maintainable as the number of systems, customers and capabilities grows. Those are architecture problems, and they are the subject of this guide.

PartContents
1. FoundationsWhat an enterprise agent is, where it creates value, maturity levels, and the design goals that drive every later decision.
2. Reference architectureLayers, deployment topologies, the lifecycle of a request, and agents that span several systems.
3. Building blocksChannels, connectors, tools, identity and permissions, tool selection, prompt architecture, context, knowledge, analytics, workflow, documents, rendering, the reasoning loop, models and multi-tenancy.
4. Operating in productionObservability and evaluation, security, cost, scaling and change management.
5. Architecture decisionsFrameworks, multi-agent designs and technology choices.
6. Build it yourselfThe open-source reference implementation, how to extend it, a production checklist and an adoption roadmap.

Part 1. Enterprise AI agent foundations

What an enterprise AI agent is

An enterprise agent is software that receives a goal in natural language, decides which operations in one or more business systems will achieve it, performs those operations on behalf of a specific person, and reports the result. Three properties distinguish it from a chatbot or a search assistant:

  • It acts. It creates, updates, submits and approves records, sends messages and produces documents, not only answers.
  • It is delegated. Every action is performed as an identified person, within that person’s permissions, and is recorded in the system’s own audit trail.
  • It is grounded. Its answers come from live records and computed results, not from what the model remembers or estimates.
CapabilityChatbot or search assistantEnterprise agent
Source of answersModel knowledge, indexed documentsLive records, computed analyses, indexed process knowledge
ActionsNoneTransactions, approvals, messages, documents
IdentityAnonymous or a shared service accountThe signed-in person, as the business system sees them
NumbersGenerated by the modelComputed by code, explained by the model
Rules and approvalsDescribed in a promptEnforced by the system of record
AccountabilityA chat transcriptThe system’s audit trail plus a structured interaction log

Where enterprise agents create value

The same architecture serves every business function. What changes is the set of connected systems and tools.

FunctionTypical systemsExamples of agent work
Sales and customerCRM, ERP sellingQualify leads, prepare quotations, convert orders, chase overdue invoices
ProcurementERP buying, supplier portalsRaise requests, compare supplier offers, track deliveries, match invoices
FinanceGeneral ledger, banking, taxExplain balances, prepare entries, analyse profit and cash, forecast results
Supply chainInventory, warehouse, manufacturingStock positions, reorder needs, slow stock, production progress
PeopleHR, payroll, attendanceLeave and expense requests, headcount and cost analysis
ServiceService desk, field serviceTriage tickets, summarise history, update status, notify customers
ManagementAll of the aboveKPI dashboards, period comparisons, exception reports, executive summaries

Maturity levels

Organisations rarely move straight to autonomous agents. Each level builds on the guarantees of the previous one, and the architecture in this guide supports all four.

LevelWhat the agent doesWhat must already be in place
1. InformedAnswers questions from live data and process knowledgeConnectors, read tools, delegated identity, rendering
2. AssistedPrepares transactions and documents for a person to confirmWrite tools, confirmation, audit, refusal handling
3. OrchestratedWalks multi-step processes and approval workflows end to endWorkflow awareness, document chains, notifications
4. AutonomousPursues standing goals on a schedule (collections, reorder, period close)Service identities, budgets, exception routing, strong evaluation

Design goals and trade-offs

Every decision in this guide serves one of six goals. They conflict in places, and the architecture is largely the record of how we resolved those conflicts.

GoalWhat it meansHow the architecture achieves it
CorrectnessEvery fact and number is trueLive reads through tools; computation in code; system refusals passed through unchanged
GovernanceThe agent never exceeds the person’s authorityDelegated identity; layered authorization; confirmation before any commit
System authorityBusiness rules live in one placeRules, validations and approvals stay in the system of record; the agent reads their outcome
Predictable costCost per request is known and boundedPer-request tool selection, layered prompts, caching, budgets
PortabilityNew systems without rewriting the agentA canonical model and a connector contract; an engine that never imports system code
EvolvabilityImprovements reach every customer safelyOne shared engine, tests for every guard, a review loop over real traffic

The most consequential of these is system authority. Early in the project it was tempting to encode business rules in prompts: discount limits, approval thresholds, mandatory fields. Rules expressed in a prompt are suggestions to a probabilistic model; they drift from the real configuration and they differ for every customer. We moved all of them back into the systems of record and designed the agent to read outcomes instead. The agent proposes; the system decides; the agent reports exactly what the system said.

Part 2. Enterprise AI agent reference architecture

Layers

The architecture has five layers. Requests flow down from the channel to the system of record; results flow back up. Each layer has a single responsibility, which keeps the parts that change often (prompts, tools, connectors) away from the parts that must stay stable (identity, authorization, execution).

Reference architecture of an enterprise agent People work in channels. Requests go to the agent runtime, which authenticates, selects tools, assembles the prompt and context, runs the reasoning loop and executes tool calls through a permission gateway and connectors into the systems of record. The runtime uses capability services for models, knowledge and analytics, and records state and logs in its own data stores. 1. Channels Chat embedded in business applications · Web and mobile · Email inbox · Document intake 2. Agent runtime Identity and sessionperson, roles, tenant Tool selectionrelevant, permitted Prompt and contextlayered, budgeted Reasoning loopsteps, guards, state Permission gatewayauthorize, confirm Rendererstable, card, chart, file 3. Capability services Language models Knowledge (RAG) Analytics engine 4. Connectors and 5. systems of record Canonical entities mapped to each system’s objects · calls run as the person ERP CRM HR and payroll Service desk Runtime data Tenants, sessions, turn state Interaction log, usage Vectors, caches
Figure 1. Reference architecture. The full platform view, with every service, is on the architecture page.
LayerResponsibilityChanges how often
ChannelsCollect the request, display structured results. No business logic.Rarely
Agent runtimeIdentity, tool selection, prompt and context, reasoning, authorization, renderingOften (prompts, guards), behind tests
Capability servicesModels, retrieval, exact analyticsIndependently, behind stable contracts
ConnectorsTranslate canonical operations into each system’s APIWhen a system or mapping changes
Systems of recordData, rules, permissions, workflowsOwned by the customer, never by the agent

Deployment topologies

The same runtime supports four ways of being deployed. The choice depends on who operates the agent and where data may travel.

TopologyHow it worksBest for
Embedded with a hosted relayA small add-in inside the business system runs the calls; the hosted runtime reasons and returns call specifications. Credentials never leave the system.SaaS delivery to many organisations with no infrastructure on their side
Self-hosted single organisationThe runtime runs in the organisation’s own environment and calls the system’s API directly.Full control, regulated environments, internal platforms
Hosted multi-systemOne runtime with several connectors serves an organisation across ERP, CRM and service systems.Cross-functional work, shared analytics
HybridReasoning hosted; execution, knowledge or analytics kept inside the organisation’s network.Data residency requirements

Noviz runs the first topology for ERPNext and Business Central, and publishes the second as open source. Both use the same engine; only the edges differ.

Lifecycle of a request

Consider the request “Which customer invoices are more than 30 days overdue, and what should we do first?”

  1. Identify. The channel sends the request with the organisation’s key and the person’s identity and roles as the business system reports them. The runtime resolves the organisation, its subscription tier and budget, and the tools this person may use.
  2. Resume or start. If the request continues a conversation, its saved turn state is loaded and the history is trimmed to a fixed budget. A new conversation may be answered from cache.
  3. Select tools. A deterministic router maps the request to business modules (here, receivables). Only the relevant, permitted tools are offered, plus a small always-available spine that includes a tool-search capability.
  4. Assemble the prompt. A stable core, the sections for the matched modules, the rule blocks of the offered tools, the organisation’s own policy, and a short session block.
  5. Add context. Conversation state and the closest knowledge passages, within budgets.
  6. Reason. The model returns either an answer or tool calls.
  7. Authorize. Each call passes the permission gateway. Reads proceed; writes become proposals for the person to confirm.
  8. Execute. The connector performs the call in the business system as the person. The system applies its own permissions and validations.
  9. Loop. Results return to the model, which decides the next step. Guards stop loops that cannot converge.
  10. Render and record. The final response is a structured object rendered identically in every channel, and the turn is written to the interaction log with tools, tokens, latency and later feedback.

Agents that span several systems

Most real work crosses system boundaries: a customer complaint in the service desk relates to an order in the ERP and a contact in the CRM. Because every connector exposes the same canonical operations, one runtime can hold several connectors and the model can combine their tools in a single turn. Three rules keep this safe:

  • Identity per system. The person is resolved separately in each system, and each call runs with that system’s identity and permissions.
  • Canonical names are unique. Entities are namespaced by system where they would collide (a CRM account and an ERP customer are different entities with an explicit link).
  • No hidden joins. Cross-system matching is done by explicit tools with explicit keys, so a result can always be traced back to the records it came from.

Part 3. Building blocks of an enterprise AI agent

Each building block below follows the same structure: its purpose, the design we use, the decisions behind it, and code from the open-source reference implementation.

Channels and embedding

People should not leave their system to use the agent. The most effective channel is a chat surface embedded in the business application itself, opened from the same menu as the rest of their work, with no second sign-in. In the embedded topology, the add-in inside the business system has a second, more important job: it executes the agent’s data calls in the person’s own session.

A turn is an exchange of structured messages between the add-in and the runtime:

  1. The add-in sends the message, the person’s identity and roles, and the turn identifier, authenticated with the organisation’s key.
  2. The runtime reasons and, instead of calling the system itself, returns call specifications: generic operations such as list, get, insert, submit, run report.
  3. The add-in executes each specification through the system’s own API, as the person, and posts the results back.
  4. The runtime resumes from saved state and continues until the final response.
// Relay to plugin: "run these calls for me, as this user"
{
  "turn_id": "6f1c0b2e-...",
  "status": "calls",
  "calls": [
    { "id": "c1", "method": "get_list", "doctype": "Sales Invoice",
      "fields": ["name", "customer", "due_date", "outstanding_amount"],
      "filters": [["docstatus", "=", 1], ["outstanding_amount", ">", 0],
                  ["due_date", "<", "2026-09-08"]],
      "order_by": "due_date asc", "limit": 50 }
  ]
}

// Plugin to relay: the results, exactly as ERPNext returned them
{ "turn_id": "6f1c0b2e-...", "results": [ { "id": "c1", "ok": true, "rows": [ ... ] } ] }

Why this design. We evaluated giving the runtime a service account with broad access. It is simpler, but it turns the runtime into a high-value target, bypasses the system’s own permissions and makes every record look as if one service user created it. Executing inside the person’s session gives native permissions and a native audit trail for free, and the organisation opens no inbound ports.

Design note

Keep the add-in thin: execute calls, display results, store settings. Every piece of reasoning placed in the add-in has to be released through each platform’s marketplace; every piece placed in the runtime reaches all organisations on the next deployment.

Connector layer and canonical model

The runtime never uses a system’s native names. It works with canonical entities such as customer, sales_order or leave_application, and canonical fields such as date, total and status. A connector translates them into the system’s own objects and fields, and implements one interface that every part of the runtime depends on.

// Every connector implements one interface (excerpt from the open-source agent)
export interface SystemConnector {
  loginWithPassword(identifier: string, password: string): Promise<UserCredential>;
  getUserRoles(identifier: string): Promise<string[]>;
  list(entityKey: string, credential: UserCredential,
       params?: { filters?: Record<string, any>; limit?: number; sortBy?: string }): Promise<any[]>;
  get(entityKey: string, credential: UserCredential, id: string): Promise<any>;
  create(entityKey: string, credential: UserCredential, data: Record<string, any>): Promise<any>;
  update(entityKey: string, credential: UserCredential, id: string, data: Record<string, any>): Promise<any>;
  submit(entityKey: string, credential: UserCredential, id: string): Promise<any>;
  aggregate(entityKey: string, credential: UserCredential, params: AggregateParams): Promise<any>;
  runReport(reportKey: string, credential: UserCredential, filters?: Record<string, any>): Promise<any[]>;
}

// The ERPNext mapping: canonical entity -> native doctype and fields
sales_order: {
  doctype: "Sales Order",
  fieldMap: { id: "name", customer: "customer", status: "status",
              total: "grand_total", date: "transaction_date" },
},
sales_invoice: {
  doctype: "Sales Invoice",
  fieldMap: { id: "name", customer: "customer", status: "status", total: "grand_total",
              outstanding_amount: "outstanding_amount", due_date: "due_date", date: "posting_date" },
},

Source: backend/src/core/types.ts, backend/src/erpnext/entityMaps/selling.ts in the open-source agent

The contract has four parts:

PartPurpose
Operationslist, get, create, update, submit, aggregate, count, run report, document output, workflow description
Entity mappingCanonical entity to native object; canonical field to native field; child tables for line items
Report mappingCanonical report keys and filters to the system’s own reports
Capability hintsWhich tools a given system or add-in can run, and any per-system shaping of a call

A product application registers its connector once at start-up, before the engine loads, and the engine reads every system through a registry. When we brought the same engine to Microsoft Dynamics 365 Business Central, the work was a new connector, its entity maps and an add-in; the reasoning, prompts, tools, gateway and renderers were unchanged.

Pitfall. Canonical names are a product decision, not a technical one. Choose them from the business vocabulary your users speak, keep them stable, and make the connector, not the prompt, responsible for every system-specific detail.

Tool design

The model can act only through tools. A tool has a unique name, a description written for the model, a JSON schema for its arguments and a handler. Tools are grouped into modules and registered centrally at start-up.

export interface ToolDefinition {
  name: string;               // "<module>.<action>", globally unique
  description: string;        // written for the model: when to use it, what it returns
  module: string;
  parameters?: object;        // JSON schema for the arguments
  handler: (args: any, session: Session) => Promise<any>;
  promptRules?: string[];     // guidance injected only when this tool is offered
}

export const knowledgeModule: MCPModule = {
  name: "knowledge",
  description: "ERP process knowledge and live approval workflow",
  tools: [ knowledgeSearchTool, workflowDescribeTool ],
};

Source: backend/src/core/types.ts, backend/src/modules/knowledge/index.ts in the open-source agent

Most tools are generated rather than written:

FactoryInputGenerated tools
Entity factoryAn entity configuration: canonical fields and allowed operations<entity>.list, .get, .create, .update, .submit
Report factoryA report configuration mapped to the system’s reportOne tool per named report
Workflow factoryA state transitionOne gated tool per transition

Hand-written modules cover what configuration cannot: analytics, charts, reports and documents, schema discovery, inbox actions, document intake and knowledge retrieval.

Granularity. We tried both extremes. One generic “query anything” tool gives the model too much freedom and produces malformed queries; hundreds of narrow tools overwhelm it. The balance that works is typed entity tools for everyday operations, a small number of powerful typed tools (aggregate, join, report) for analysis, and discovery for everything else.

Descriptions are part of the product. A description that states limits and alternatives (“returns at most 50 rows; use analytics.aggregate for totals”) prevents more errors than any prompt instruction, because the model reads it at the exact moment it chooses the tool.

Identity, permissions and governance

Authorization is layered. Each layer can only remove options; none can add access that another layer denied.

  1. Organisation. The key resolves one organisation, its active term, its tier (read-only or read and write) and its budget.
  2. Person. The signed-in person and their roles, as reported by the business system.
  3. Role allowance. Roles resolve to the tools that person may use. A sales user is never offered payroll tools.
  4. Tool selection. Only relevant permitted tools are offered for this request; discovery never reaches beyond the grant.
  5. Gateway. Every call, from the model or from any other route, passes one function. Unknown or ungranted tools are rejected; writes become proposals.
  6. System of record. The call runs as the person, so the system applies its own permissions, validations and approval workflow.
// The single enforcement and execution point for every tool call
export async function callTool(session: Session, toolName: string, args: any) {
  const allowed = session.allowed_tools.includes("*") || session.allowed_tools.includes(toolName);
  if (!allowed) throw new ToolNotAllowedError(toolName);

  const tool = moduleRegistry.findTool(toolName);
  if (!tool) throw new Error(`Unknown tool: ${toolName}`);

  return tool.handler(args, session);   // the handler calls the connector as session.credential
}

Source: backend/src/core/gateway.ts (business-rule hooks omitted) in the open-source agent

A single entry point makes the permission model auditable: to verify it, you read one function. Confirmation is part of the same design. Any tool that creates, updates, submits, cancels or sends returns the exact document and values it intends to write; the person confirms in the interface; only then is the call executed.

Important

The agent must never hold a role of its own. If it can do something the person cannot, the permission model has a hole that no prompt can close.

Tool selection at scale

An agent for a full business system has several hundred tools. Offering all of them on every request is slow, expensive and less accurate, and model providers cap the number of tools per request. We select tools per request in three stages: role allowance, a deterministic module router, and a spine with discovery.

const RELAY_SPINE_TOOLS = new Set([
  "tools.search",
  "data_table.list",
  "data_table.search_schema",
  "database_engine.execute_query",
  "report.find",
  "analytics.catalog",
]);

export function selectRelayTools(tools: ToolDefinition[], prompt: string): ToolDefinition[] {
  const modules = relayModulesFor(prompt);
  const wantsCharts = detectExplicitModules(prompt).has("analytics");
  return tools.filter((t) => {
    if (RELAY_SPINE_TOOLS.has(t.name) || t.module === "context") return true;
    if (wantsCharts && (t.name === "analytics.aggregate" || t.name === "chart.build")) return true;
    if (modules.size > 0 && (t.name.endsWith(".list") || t.name.endsWith(".get")) && toolMatchesBusinessModule(t, modules)) {
      return true;
    }
    return false;
  });
}

Source: backend/src/core/toolRelevanceFilter.ts in the open-source agent

The router matches the request against keyword lists per business module and widens related pairs (a customer question about invoices also loads selling, because that is where invoices live). The same router decides which prompt sections are loaded, so the tools offered and the guidance given can never disagree. Anything not offered remains reachable through tools.search; tools found that way stay available for the rest of the turn.

Why deterministic. We also built model-based tool selection, an extra model call that picks tools. It is useful for very broad roles such as system administrators and is kept for them. For department roles, the keyword router is faster, free, predictable and easy to test.

Prompt architecture

A single large system prompt is the most common source of cost and inconsistency in agent projects: every rule is paid for on every request, and rules for one domain interfere with another. We assemble the prompt from layers.

LayerContentWhen included
CoreIdentity, operating principles, tool guidance, display conventionsAlways; identical bytes on every request
Write layerHow to propose changes and wait for confirmationOnly when the person may write
Module sectionsDomain guidance per business moduleOnly for modules the router matched
Tool rule blocksGuidance a specific tool needsWhen that tool is offered, once per conversation
Organisation policyThe customer’s own instructionsNew conversations, that organisation only
Session blockDate, company, person, rolesAlways; always last
export function buildSystemPrompt(prompt: string, canWrite: boolean, identity: AppIdentity, frappeUser: string, frappeRoles: string[]): string {
  const realModules = [...relayModulesFor(prompt)].filter((m) => BUSINESS_MODULE_KEYS.includes(m));

  const sections: string[] = [];

  if (realModules.length === 0) {
    sections.push(CORE_SYSTEM_PROMPT);
  } else {
    sections.push(THIN_CORE, AUTOFLOW_RULES); // AUTOFLOW: any real business-data turn wants the document-chain facts
  }

  if (canWrite) sections.push(WRITE_OPERATIONS);

  const seen = new Set<string>();
  for (const m of realModules) {
    for (const section of MODULE_PROMPT_SECTIONS[m] || []) {
      if (!seen.has(section)) {
        seen.add(section);
        sections.push(section);
      }
    }
  }

  sections.push(buildUserContext(identity, frappeUser, frappeRoles));
  return sections.filter(Boolean).join("\n\n");
}

Source: backend/src/systemPrompt/index.ts in the open-source agent

The order is fixed for a reason. Model providers cache the beginning of a prompt that is byte-for-byte identical across requests and charge less for it. Keeping the stable core first and the per-person session block last lets most of every request hit that cache.

Where guidance goes

The open-source agent ships every prompt block empty, so each organisation writes its own. The examples below show the shape only.

GuidanceLocationExample
Who the agent issystemPrompt/core/identity.ts"You are the operations agent for Acme."
How to choose toolssystemPrompt/core/toolDiscovery.ts"For totals, use analytics.aggregate."
How to present resultssystemPrompt/core/display.ts"Show lists as tables."
Before any changesystemPrompt/core/writeOperations.ts"Show the values and wait for confirmation."
One business areasystemPrompt/modules/<module>.ts"A quotation becomes a sales order."
When a module loadsMODULE_KEYWORDS in systemPrompt/modules/index.tsselling: ["quotation", "sales order"]
One toolThe tool’s promptRulespromptRules: [REPORT_RULES]
Per-message hintssystemPrompt/core/hints.ts"Use a relative date filter."
Organisation policiesAdministration console (no code)A policy document
// systemPrompt/modules/selling.ts
export const SELLING_MODULE = "A quotation becomes a sales order.";

// systemPrompt/modules/index.ts
selling: ["quotation", "sales order"],

Placeholder wording. A request that mentions “quotation” loads the selling section and the selling tools together.

export const REPORT_RULES = "Use reports for files, not chat tables.";

{ name: "report.generate", /* ... */ promptRules: [REPORT_RULES] }

Placeholder wording. The rule is sent once per conversation, the first time the tool appears.

Important

Treat prompts as production code: version them, review them, and run the full regression suite on every change. A rule added to fix one conversation can quietly break ten others.

Context and memory

Context is everything the model sees besides the request and the tools. We organise it in three tiers by cost.

TierSourceUse
HotThe current conversation, recently referenced records, the last result setAlways included, within a budget
WarmKnowledge and policy passages closest to the requestAdded automatically, within a budget
ColdEverything else in the systems and the knowledge storeFetched only through tools, when the model asks
  • Budgets per source stop one large source from crowding out the others.
  • Bounded history. Conversations keep a fixed number of earlier turns. We drop older turns rather than ask the model to summarise them, because silent loss of detail in a summary is harder to detect than a clear limit.
  • Large results stay on the server. The model receives a page and a summary; the full set remains available for paging, export and analysis.
  • Memory is scoped to the session, not the person, so two sign-ins never inherit each other’s state.

Knowledge and retrieval (RAG)

Tools give the agent access to data. They do not tell it what the data means: that a delivery reduces stock, that a submitted invoice can no longer be edited, that a purchase order in “Pending Approval” is waiting for a specific role. Retrieval-augmented generation supplies that meaning from documents at the moment it is needed, without carrying it in every prompt.

SourceExample contentKept current by
Process documentationOrder to cash, procure to pay, stock movements, period close, request and approval flows, document lifecycleVersioned files, re-indexed on change
Organisation documentsPolicies, standard operating procedures, contractsRe-indexed on every edit; can be deactivated
Live workflow definitionsThe organisation’s own approval workflows: states, transitions, rolesRead from the system at request time, never copied

Ingestion

  1. Parse each document and its metadata: title, module, the document types it describes.
  2. Chunk by meaning at section headings, splitting long sections with a small overlap so a sentence on a boundary keeps its context.
  3. Label each chunk “document title > section” so it remains understandable on its own.
  4. Embed and store in PostgreSQL with pgvector, linked to the source document; re-indexing replaces only that document’s chunks.
  5. Skip unchanged sources by checksum, so embeddings are paid for once.
/** Splits a document into labelled retrieval chunks: one per "## " section,
 *  long sections further split by the shared policy chunker. */
export function chunkKnowledge(doc: KnowledgeFile): KnowledgeChunk[] {
  const sections: { name: string; text: string }[] = [];
  let current = { name: "Overview", text: "" };
  for (const line of doc.body.split("\n")) {
    const h2 = line.match(/^##\s+(.+)$/);
    if (h2) {
      if (current.text.trim()) sections.push(current);
      current = { name: h2[1].trim(), text: "" };
    } else if (!/^#\s+/.test(line)) {
      current.text += line + "\n";
    }
  }
  if (current.text.trim()) sections.push(current);

  const chunks: KnowledgeChunk[] = [];
  for (const s of sections) {
    const parts = chunkText(s.text.trim());
    parts.forEach((part, i) => {
      const label = `knowledge:${doc.slug}#${s.name}${parts.length > 1 ? ` (${i + 1}/${parts.length})` : ""}`;
      chunks.push({ label, content: `${doc.title} > ${s.name}\n\n${part}` });
    });
  }
  return chunks;
}

Source: backend/src/core/knowledgeStore.ts in the open-source agent

Retrieval

Knowledge reaches the model in two ways: automatically, as the closest passages in the warm context tier, and on demand through a knowledge.search tool filtered by module, whose results carry their source titles for citation.

select d.title, e.label, e.content, 1 - (e.embedding <=> $1) as score
from   context_embeddings e
join   knowledge_documents d on d.id = e.knowledge_document_id
where  d.active and ($2::text is null or d.module = $2 or d.module is null)
order  by e.embedding <=> $1
limit  $3;

Source: backend/src/core/knowledgeStore.ts (search) in the open-source agent

Workflow awareness

Static documents describe the standard process; the organisation’s real approval process lives in the system and changes over time. A second tool, workflow.describe, reads it live: the configured states and transitions and, for one document, its current state and the actions the system allows this person now.

{
  "doctype": "Purchase Order",
  "has_workflow": true,
  "states": [ { "state": "Pending Approval", "doc_status": "Draft" },
              { "state": "Approved", "doc_status": "Submitted" } ],
  "transitions": [ { "from": "Pending Approval", "action": "Approve", "to": "Approved",
                     "allowed_role": "Purchase Manager", "condition": null, "self_approval": false } ],
  "document": { "name": "PUR-ORD-2026-00087", "current_state": "Pending Approval",
                "docstatus": 0, "actions_for_you": [] }
}

Output of workflow.describe (backend/src/erpnext/workflow.ts). The workflow definition is metadata; the document and the actions available on it are read as the user, through ERPNext’s own get_transitions.

Benefit. With both, the agent answers “why can’t I submit this order?” correctly: the knowledge base explains approval workflows; the live workflow shows that this order needs a role the person does not hold. No rule inside the agent made that decision.

Important

Do not use retrieval for business data or business rules. Data read from vectors is stale and ignores permissions; rules belong in the system. Retrieval is for meaning.

The analytics engine: a bridge between LLM analysis and local computation

Analysis is where enterprise agents most often fail and where they can create the most value. Managers do not only ask for records; they ask how the business is performing, where money is being lost, what next quarter looks like and which customers, products or suppliers need attention. A language model alone cannot answer these questions reliably: it is poor at arithmetic over many rows, it cannot read thousands of records within a prompt budget, and anything it estimates cannot be reproduced. A database query alone cannot answer them either: it does not understand the question, does not know which analysis fits, and cannot explain the result.

We therefore built analytics as a dedicated engine with two halves and a bridge between them. The LLM side understands the question, chooses the analysis and explains the result. The local analytics engine, running on our own servers next to the agent runtime, fetches the complete data and computes every number, comparison, forecast and anomaly exactly. The bridge is a small set of typed tools and a fixed data contract, so the two halves exchange intent and results, never raw rows.

The analytics bridge On the left, LLM analysis understands the question, chooses an analysis from the catalogue and explains the result. In the middle, the bridge carries an analysis key and parameters in one direction and a compact summary in the other. On the right, the local analytics engine fetches the data as the user, maps it to the data contract, computes the analysis in Python and renders a dashboard and a PDF. LLM analysis Understand the questionintent, period, scope Choose the analysisfrom 100 in the catalogue Explain the resultfindings, risks, next steps Follow updrill down, compare, act The bridge Typed analytics toolscatalog, build, compare Analysis key + parametersno data in the request Data contractfixed columns per dataset Summary + KPIs backa few hundred tokens Local analytics engine Fetch as the usercomplete period, paged Compute in Pythonpandas, NumPy, SciPy Forecast and detectranges, anomalies, shifts Insight and renderKPIs, dashboard, PDF
Figure 2. The analytics bridge. Intent travels right; computed results travel left; rows never cross to the model.

What the engine is for

The engine carries a catalogue of 100 standard analyses in 12 categories. Each one is a complete, tested computation, not a template for the model to fill in.

AreaQuestions it answersExamples of analyses
Financial resultsAre we profitable, where is margin made or lost, how will next periods close?Profit and loss, gross and net profit, profit margin, budget against actual, revenue against target, with forecasts
Cash and working capitalHow much cash will we have, who owes us, whom do we owe?Cash position, inflow and outflow, working capital, receivables and payables ageing, payment history
Stock flowWhat do we hold, what moves, what is dead, what must be reordered?Stock balance and valuation, turnover, stock ageing, dead stock, slow and fast movers, reorder analysis, ABC classes
Customers and salesWho is growing, who is slipping, where is revenue concentrated?Sales by customer, item, territory and salesperson, growth trends, profitability by customer, purchase frequency
Suppliers and purchasingWho delivers late, who is getting more expensive, where are we dependent?Supplier delivery performance, price comparison and trends, spend concentration, open orders
Operations and peopleAre we efficient, where is waste, what does the workforce cost?Production efficiency and wastage, workstation use, project budget and progress, payroll cost, turnover
ExecutiveHow healthy is the business overall, and what needs attention first?Business KPI dashboard, overall business health, executive summary with watch-outs

Every analysis can add a forecast for the next periods, flag anomalous periods and locate the point where a trend changed, and any analysis can be compared across two periods or two segments. This is what lets the agent address complex business issues: not just “sales fell”, but which customers drove the fall, when it started, whether it is unusual, and what the next quarter is likely to look like if nothing changes.

How a request flows through the engine

  1. Understand. The model reads the question and identifies the intent, period and scope.
  2. Choose. It calls the catalogue tool to find the matching analyses by name and purpose, and asks the person when the request is ambiguous.
  3. Request. It calls the build tool (or the period or segment comparison tool) with an analysis key and parameters only.
  4. Fetch. The runtime retrieves each dataset the analysis needs, for the complete period, through the connector and as the person. The system’s permissions decide what is visible.
  5. Normalise. Rows are mapped to the data contract: fixed columns per dataset (date, party, item, amount, quantity, account). Every system maps to the same contract.
  6. Compute. The engine runs the analysis in Python: aggregation, grouping, period arithmetic, ageing, ranking, concentration, ABC classification.
  7. Forecast and detect. A forecast blends a damped-trend (Holt) model with linear regression and reports an 80% range; anomalies are found with a robust z-score based on the median absolute deviation; a change point is located with a Welch t-test across candidate splits.
  8. Interpret. An insight layer adds KPIs with change against the prior period, concentration and trend findings, watch-outs and recommendations.
  9. Render. Charts are drawn by the engine, assembled into a dashboard for the chat and a PDF for sharing.
  10. Explain. Only the summary and KPIs return to the model, which explains them, answers follow-up questions and proposes the next analysis or action.

Why the work is divided this way

PropertyLLM aloneLLM + local analytics engine
AccuracyNumbers estimated from a sample of rowsEvery number computed in code over the complete data
ReproducibilityDifferent answer on a different runSame data and parameters give the same result
ScaleLimited by the prompt sizeLimited by the server, with chunked fetching
CostThousands of rows of tokens per questionA few hundred tokens of summary
PermissionsDepends on what was pasted into the promptData fetched as the person; nothing they cannot see
ExplanationFluent but unverifiableFluent and grounded in computed results

The engine runs as its own service because it has a different language, a different load profile and a different failure mode from the reasoning runtime: heavy numerical work should never slow down a conversation, and a failed analysis must never take a conversation down with it. It listens only on a private interface, holds no system credentials and never contacts a customer system; the runtime brings it the data.

In the open-source agent

The reference implementation ships the same pattern at a smaller scale: one analysis, a monthly trend with growth and a projection, run as a local Python process.

"""Monthly trend: the sample analysis of the Python analytics engine.

Input  (stdin, JSON): {"rows": [{"date": "2026-01-14", "total": 1200.0}, ...]}
Output (stdout, JSON): KPIs, a chart specification and observations.

The rows use canonical field names (date, total), the data contract every
connector maps its native fields to, so the analysis works for any ERP.
The rows never reach the language model; only this output does.
"""
import json
import sys

import numpy as np
import pandas as pd


def analyse(rows):
    df = pd.DataFrame(rows)
    if df.empty or "date" not in df or "total" not in df:
        return {"kpis": [], "chart": None, "observations": ["No records in the period."]}

    df["date"] = pd.to_datetime(df["date"])
    df["total"] = pd.to_numeric(df["total"], errors="coerce").fillna(0.0)
    monthly = df.set_index("date")["total"].resample("MS").sum()
    growth = monthly.pct_change().mul(100).round(1)

    # Linear projection for the next month (least squares over the series).
    if len(monthly) > 1:
        slope, intercept = np.polyfit(np.arange(len(monthly)), monthly.values, 1)
    else:
        slope, intercept = 0.0, float(monthly.iloc[0])
    projection = max(0.0, float(slope * len(monthly) + intercept))

    best = monthly.idxmax()
    observations = [
        f"Highest month: {best:%B %Y} at {monthly.max():,.2f}.",
        f"The trend is {'rising' if slope > 0 else 'falling' if slope < 0 else 'flat'} by about {abs(slope):,.2f} per month.",
    ]
    if growth.notna().any():
        observations.append(f"Latest month-on-month change: {growth.iloc[-1]:+.1f}%.")

    return {
        "kpis": [
            {"label": "Total", "value": round(float(monthly.sum()), 2)},
            {"label": "Average per month", "value": round(float(monthly.mean()), 2)},
            {"label": "Projected next month", "value": round(projection, 2)},
        ],
        "chart": {
            "type": "line",
            "labels": [d.strftime("%b %Y") for d in monthly.index],
            "series": [{"name": "Total", "values": [round(float(v), 2) for v in monthly.values]}],
        },
        "observations": observations,
    }


if __name__ == "__main__":
    print(json.dumps(analyse(json.load(sys.stdin).get("rows", []))))

Source: backend/analytics/monthly_trend.py in the open-source agent

handler: async (args, session) => {
  if (!TREND_ENTITIES.includes(args.entity)) throw new Error(`entity must be one of: ${TREND_ENTITIES.join(", ")}`);
  const rows = await systemConnector.list(args.entity, session.credential, {
    filters: { date: ["between", [args.from_date, args.to_date]], status: ["not in", ["Draft", "Cancelled"]] },
    limit: ROW_CAP,
    sortBy: "date",
    sortDir: "asc",
  });
  const result = await runPythonAnalysis("monthly_trend", rows.map((r) => ({ date: r.date, total: r.total })));
  return {
    entity: args.entity,
    period: { from: args.from_date, to: args.to_date },
    records_analysed: rows.length,
    ...(rows.length >= ROW_CAP ? { note: `Only the first ${ROW_CAP} records were analysed; narrow the date range for a complete result.` } : {}),
    ...result,
  };
},

Source: backend/src/modules/trends/index.ts in the open-source agent

import { spawn } from "child_process";
import path from "path";

/**
 * Runs one Python analysis (backend/analytics/<name>.py) as a child
 * process: rows in on stdin, JSON result out on stdout. Keeps the agent a
 * single deployment; a hosted platform can run the same scripts as a
 * separate Python service without changing them. The rows go to Python,
 * never to the language model; the model only sees the returned summary.
 */
export const ANALYTICS_DIR = path.resolve(__dirname, "../../analytics");

export function runPythonAnalysis(
  name: string,
  rows: object[],
  opts: { command?: string; args?: string[]; timeoutMs?: number } = {}
): Promise<any> {
  if (!/^[a-z0-9_]+$/.test(name)) return Promise.reject(new Error(`Invalid analysis name "${name}"`));
  const command = opts.command ?? (process.env.PYTHON_BIN || "python3");
  const args = opts.args ?? [path.join(ANALYTICS_DIR, `${name}.py`)];

  return new Promise((resolve, reject) => {
    const child = spawn(command, args, { timeout: opts.timeoutMs ?? 30000 });
    let out = "";
    let err = "";
    child.stdout.on("data", (d) => (out += d));
    child.stderr.on("data", (d) => (err += d));
    child.on("error", (e) => reject(new Error(`Could not start the analytics engine (${command}): ${e.message}`)));
    child.on("close", (code) => {
      if (code !== 0) return reject(new Error(`Analysis "${name}" failed: ${err.trim() || `exit code ${code}`}`));
      try {
        resolve(JSON.parse(out));
      } catch {
        reject(new Error(`Analysis "${name}" returned invalid JSON`));
      }
    });
    child.stdin.end(JSON.stringify({ rows }));
  });
}

Source: backend/src/core/pythonAnalysis.ts in the open-source agent

To add an analysis, write a script that reads rows in the data contract and prints KPIs, a chart specification and observations, then expose it with a tool like analytics.monthly_trend. The model never needs to know how it is computed, only what it is for.

Workflow and approvals

Business work is a chain of documents: a request becomes an order, an order becomes a delivery, a delivery becomes an invoice, an invoice is paid. Systems of record track that chain through links between documents and enforce approval workflows at specific steps. The agent’s job is to walk the chain faster, never to shorten it.

  1. Read the state of the current document and its workflow, live, as the person.
  2. Propose the next document, created from its source so the system keeps links and progress tracking.
  3. Confirm with the person, showing the exact values.
  4. Let the system validate and apply its approval workflow; pass through any refusal unchanged.
  5. Leave approvals to approvers. An approver can act in the system or through the agent if they hold the role; the agent never approves on someone’s behalf.
  6. Report and offer the next step, with a notification where the workflow expects one.

Document intake

Incoming documents (supplier invoices, customer orders, receipts) are a natural task for an agent:

  1. The person uploads a file or takes a photo in the channel.
  2. A vision-capable model extracts text and structure: party, dates, line items, totals.
  3. The agent identifies the transaction type and matches the party and items in the system.
  4. It prepares a draft and presents it as a proposal; nothing is written until the person confirms.

Rendering and experience

The model does not produce HTML and does not draw charts. It chooses a response type, and the runtime produces a structured object that every channel renders with the same components.

TypeUsed forProduced by
TextExplanations, confirmations, system refusalsThe model
TableLists of records with paging and links into the systemTable renderer
CardsOne record, an email or a notification with actionsCards renderer
Chart and dashboardTrends, breakdowns, KPI viewsChart builder or the analytics engine
FileReports, the system’s own print formats, exportsReport generator or the system
ProposalA write awaiting confirmationThe gateway
{
  "type": "table",
  "title": "Overdue customer invoices",
  "columns": ["Invoice", "Customer", "Due date", "Outstanding"],
  "rows": [["ACC-SINV-2026-00412", "Gulf Traders LLC", "2026-08-14", 18450.00]],
  "links": { "Invoice": "/app/sales-invoice/{value}" },
  "page": { "offset": 0, "size": 20, "total": 37 }
}

Benefit. One response format means one design on every surface, whether the agent is embedded in ERPNext, in Business Central or in a web app, and the model cannot alter a number on its way to the screen because tables and charts are drawn from tool results.

Reasoning loop, guards and workers

The reasoning loop is a plain loop: select the tools for this step, inject their rule blocks, call the model, check the proposed calls, execute them, append the results, repeat. Tools found through discovery stay available for the rest of the turn.

const toolsForThisStep = () => {
  const selected = selectRelayTools(allowedTools, relevanceSearchText);
  const selectedNames = new Set(selected.map((t) => t.name));
  const forced = allowedTools.filter((t) => forcedToolNames.has(t.name) && !selectedNames.has(t.name));
  return [...forced, ...selected].slice(0, MAX_TOOLS_PER_REQUEST);
};

for (let i = 0; i < appConfig.llm.maxToolIterations; i++) {
  const tools = toolsForThisStep();
  for (const t of tools) for (const block of t.promptRules ?? []) injectPromptBlock(messages, block);
  const response = await this.llm.chat(messages, tools);
  if (response.tool_calls.length === 0) { finalText = response.content || ""; break; }
  messages.push({ role: "assistant", content: response.content || "", tool_calls: response.tool_calls });
  for (const call of response.tool_calls) {
    // guards: repeated calls, denied tools, step and time budgets ...
    const result = await callTool(session, call.name, call.arguments);
    if (call.name === "tools.search") {
      for (const name of foundToolNames(result).slice(0, TOOLS_SEARCH_FORCE_CAP)) forcedToolNames.add(name);
    }
    messages.push({ role: "tool", content: JSON.stringify(result), tool_call_id: call.id, name: call.name });
  }
}

Source: backend/src/core/reasoningEngine.ts (shortened) in the open-source agent

Guards. Models sometimes repeat a call that cannot succeed, such as paging through thousands of rows to build a total. Guards detect these patterns and redirect the model to the right tool, an aggregate or a report. Each guard exists because of a reproduced failure and has a test that names it; we remove a guard only when that failure is proven impossible.

State and workers. The runtime is stateless between steps: conversation and turn state live in the database and cache. Several workers serve requests behind one endpoint and any worker can continue any turn. Long analyses and reports can move to a background queue so that web requests return quickly.

Why plain functions. We kept the loop as ordinary, named functions rather than a graph or framework abstraction, so that each step can be read, tested and traced on its own, and so that a guard can sit exactly between “the model proposed” and “the tool executed”.

Model connectivity and strategy

The runtime talks to models through one provider interface. Any service implementing the chat-completions format with tool calling can be used.

export interface LLMProvider {
  chat(request: {
    messages: ChatMessage[];
    tools?: ToolSchema[];
    temperature?: number;
    maxTokens?: number;
  }): Promise<{ content: string | null; toolCalls: ToolCall[]; usage: TokenUsage }>;
}

Illustrative. The open-source implementation is backend/src/providers/llm/openaiProvider.ts

  • Switchable at runtime. Base URL, model and key are settings; changing provider takes seconds and no deployment.
  • One model per task. Reasoning, document vision and embeddings can use different providers, so one provider’s limit or outage does not stop everything.
  • Timeouts, retries and clear failures on every call.
  • Rate and budget gates. A shared token-rate gate keeps all workers under provider limits; a monthly budget per organisation bounds cost.
  • Usage recorded per request: prompt, cached and completion tokens, which is how cost problems are found and fixed.

Multi-tenancy and customization

A hosted agent serves many organisations from one runtime. Four mechanisms keep them separate and individual:

MechanismWhat it isolates or adapts
Organisation keys and tiersWho is calling, whether they may write, their term and budget
Scoped dataTurn state, logs, usage, knowledge and caches are keyed by organisation
Organisation policyPlain-language instructions that apply only to that organisation
Customization layersStandard behaviour, then a system layer, then a client layer that can extend, override or add, and switch modules on or off; the base is never edited

Benefit. One engine and one release serve every organisation, yet each can have its own behaviour without a fork, the same way ERPNext and Business Central themselves are customised.

Part 4. Running an enterprise AI agent in production

Observability and evaluation

An agent is only as reliable as the evidence about how it behaves. We rely on four instruments:

  • Interaction log. Every request records the prompt, the modules matched, the tools offered and called, the gates that fired, the response type, latency, token usage and later user feedback. It is the single source for every improvement.
  • Review loop. Negative feedback, blocked gates and slow or failed turns are reviewed across all organisations. A pattern found in one organisation’s traffic is fixed once, centrally, and every organisation receives the fix on the next release.
  • Automated tests. The engine has more than 1,300 automated tests covering tool generation, permissions, prompt assembly, tool selection, guards, renderers and connector mapping. Every change runs the full suite.
  • Regression discipline. Before a guard, prompt layer or session behaviour is changed, its history is read and the change is shown not to reopen the failure it was added for.

Security and compliance

ConcernControl
Excess privilegeDelegated identity; layered authorization; no agent role
Unwanted changesConfirmation before every commit; read-only tier
Credential exposureEmbedded execution: credentials never leave the system. Self-hosted: per-person credentials encrypted at rest
Prompt injectionContent from records, emails and documents is treated as data, never as instructions; it cannot grant tools, change permissions or skip confirmation, because those are enforced in code
Data leakage between customersEvery store and cache keyed by organisation; retrieval filtered by organisation
Data minimisationLarge result sets and analytics rows stay on the server; the model receives pages and summaries
Mailbox safetyMail is read without deletion or moving
ExposureServices listen privately behind one gateway; no inbound access to customer systems
AccountabilitySystem audit trail as the real person; interaction log per request

Cost management

Cost per request is an architectural property. The largest savings in our platform came from design, not from cheaper models:

  • Tool selection sends a few dozen tool schemas instead of hundreds.
  • Layered prompts send only the guidance a request needs, with a stable prefix that providers cache.
  • The analytics bridge replaces thousands of rows with a summary of a few hundred tokens.
  • Response caching answers repeated first questions exactly and very similar ones semantically, scoped per organisation.
  • Budgets and gates make the worst case bounded and visible.

Scaling

  • Modular monolith first. The runtime is one service with clear modules. A part becomes its own service only when it has a different language, load or failure profile: analytics and knowledge qualify; the reasoning hot path does not, because a network hop between deciding and acting adds latency and drift.
  • Stateless workers behind one endpoint, with state in PostgreSQL and Redis.
  • Background turns for long analyses and reports.
  • Provider capacity monitored per minute, with headroom planned before limits are reached.

Change management

  • One source of truth. What runs in production is exactly what is in the main branch.
  • Tested deployments. Tests first, build beside the running version, swap, health check, automatic rollback, nothing left behind.
  • Prompts and guards as code, reviewed and tested like any other change.
  • Narrow changes. Fix the root cause in the smallest part that owns it; avoid patching symptoms in the prompt.

Part 5. Architecture decisions

Frameworks: why we built our own engine

LangChain, LangGraph, CrewAI, AutoGen and Semantic Kernel are valuable for prototypes and for applications that mainly connect many third-party sources. We evaluated them and chose a small engine of our own, for reasons specific to agents that act on systems of record.

ConcernWith a general frameworkWith our own engine
Prompt controlPrompts assembled by templates; exact bytes hard to see and stable across versionsEvery byte and its order controlled, enabling prompt caching and reproducible behaviour
Token costAll tools and long default instructionsPer-request tool selection and layered prompts
PermissionsAdded around the frameworkOne gateway is the only path to a tool
GuardsHard to place between proposal and executionNamed, tested functions in the loop
DebuggingFailures deep inside framework stacksA short loop; each step visible in the log
DependenciesFrequent breaking changes, large dependency treeA few mature libraries that change rarely
CapacityFrameworks organise code; they do not add capacityCapacity from workers, caching and provider limits we manage directly

If you start with a framework, keep the gateway, the connector contract and the tool definitions in your own code. Then orchestration can be replaced later without rebuilding the system.

Multi-agent designs

Several agents talking to each other multiply token use and latency, and make failures harder to trace. For conversational work in business systems, one agent with many well-selected tools performs better. Specialist agents earn their place when they have standing goals of their own, such as collections, replenishment or period close, or when work runs in parallel for a long time. We build them as instances of the same engine with their own tool set and prompt layer, behind the same gates, rather than as a separate agent framework.

Technology choices

AreaChoiceReason
RuntimeNode.js, TypeScript, ExpressTyped tool schemas and connectors; strong concurrency for API-bound work
Data and vectorsPostgreSQL with pgvectorOne database for state, logs, settings and vectors
Cache and coordinationRedisResponse cache, rate limits, token-rate gate across workers
AnalyticsPython: pandas, NumPy, SciPy, matplotlibThe standard numerical stack, run as a private service
DocumentsPDF generation and text extraction librariesReports out; uploaded PDFs and Word files in
InterfacesEmbedded pages in each system; React for webOne response format rendered natively
TestsJest; pytest for analyticsFast tests for every component

Part 6. Build your own enterprise AI agent

The open-source reference implementation

The Noviz agent is published as open source under AGPL-3.0 as a self-hosted reference implementation of this architecture. It uses ERPNext as its system of record and includes the connector contract and ERPNext connector, the tool factories and registry, the gateway and role-based permissions, per-request tool selection, the layered prompt pipeline (shipped empty for your own guidance), context tiers, knowledge retrieval with live workflow awareness, a Python analytics sample, renderers, reports, document intake and the agent and administration web apps.

github.com/Zeetech-soft-solution/erp-ai-agent

Prerequisites

  • A system of record with an API (the reference uses ERPNext 15 or 16 and an API key for a service user)
  • Node.js 20 or later, PostgreSQL 15 or later with the vector extension, Python 3 with pandas and NumPy
  • A key for an OpenAI-compatible model provider

Install

git clone https://github.com/Zeetech-soft-solution/erp-ai-agent.git
cd erp-ai-agent/backend
npm install
cp .env.example .env          # ERPNext URL and key, LLM key, DATABASE_URL, CREDENTIAL_ENCRYPTION_KEY

# Create the schema: run every migration in filename order
for f in src/db/migrations/*.sql; do psql "$DATABASE_URL" -f "$f"; done

npm test                      # the full test suite

# Optional: knowledge base and Python analytics
npm run knowledge:index                          # embed backend/knowledge/**/*.md
pip install -r analytics/requirements.txt        # pandas and numpy for analytics/*.py

npm run dev                   # API on the port set in .env

cd ../frontend/agent && npm install && cp .env.example .env && npm run dev
cd ../admin && npm install && cp .env.example .env && npm run dev

Code structure and why it is organised this way

backend/src/
  core/            engine: gateway, reasoning loop, module registry, factories,
                   context assembly, tool filtering, renderers registry, stores
  config/          canonical entities, role policy, reports, workflows, settings
  erpnext/         the ERPNext connector: client, entity maps, report map
  modules/         hand-written tools: analytics, trends, chart, reports, documents,
                   schema, tool discovery, inbox actions, context, knowledge
  providers/       LLM, embeddings, vision, context tiers
  renderers/       table, cards, chart
  systemPrompt/    layered prompt: core blocks and module sections
  routes/          auth, agent, admin, tools, policy documents, webhooks
  db/migrations/   PostgreSQL schema
backend/knowledge/ process documents for the knowledge base (RAG)
backend/analytics/ Python analyses, run as a child process
frontend/
  agent/           the user-facing agent app (React)
  admin/           the administration console (React)

Each area of the code exists for a reason, and the boundaries between them are what make the agent portable and maintainable.

AreaPurposeBenefit of keeping it separate
core/The engine: gateway, reasoning loop, registry, factories, context, tool selection, storesNever imports a system’s code, so it works unchanged for any system of record
core/gateway.tsThe single path to every toolPermissions can be audited by reading one function
config/Canonical entities, role policy, reports, workflows, settingsNew coverage is configuration, not code, and is reviewed as data
erpnext/The connector: client, entity maps, report map, workflow readerEverything system-specific in one folder; a second system is a sibling folder
modules/Hand-written tools: analytics, trends, charts, reports, documents, schema, discovery, inbox, knowledgeEach capability can be added or switched off on its own
systemPrompt/Layered prompt: core blocks, module sections, hints, keywordsGuidance changes without touching logic; loaded only when relevant
providers/Models, embeddings, vision, context tiersA provider can be replaced by writing one class
renderers/Table, cards, chartOne output format for every channel
routes/HTTP surface: sign-in, agent, admin, tools, knowledgeTransport kept away from reasoning
db/migrations/The database schema, in orderReproducible environments
knowledge/Process documents for retrievalMeaning maintained as documents, not prompt text
analytics/Python analyses run by the engineExact computation in the right language, outside the model
frontend/Agent and administration appsInterfaces that only render structured responses

Extend the agent

TaskWhat to change
Add an entityAdd its configuration (canonical fields and operations) and its mapping in the connector. The factory generates the tools.
Grant it to a roleAdd the tool names to that role in the role policy.
Add a custom toolWrite a module with a definition and handler and register it; the gateway applies permissions automatically.
Connect another systemImplement the connector interface and entity maps for it (a CRM, HR or service desk); tools, prompts, gateway and interfaces stay unchanged.
Add process knowledgeAdd a Markdown document with title, module and sections under knowledge/ and index it.
Add an analysisAdd a script under analytics/ that reads the data contract and returns KPIs, a chart and observations; expose it with a tool.
Write guidanceFill the prompt blocks described in Where guidance goes.
Change model providerSet base URL, model and key.

Production checklist

  • Every tool call passes one gateway; no handler is called directly.
  • Every system call runs as the signed-in person; a service identity is used only for metadata.
  • Writes require explicit confirmation; read-only access is available.
  • Business rules live in the system of record; refusals are shown as returned.
  • Numbers and analyses are computed by code; rows are not sent to the model.
  • Tools are selected per request; prompts are layered with a stable prefix.
  • Knowledge explains processes; live data and workflows are read through tools.
  • Content from records, emails and files is treated as data, never as instructions.
  • Timeouts, retries, rate limits and budgets are configured.
  • Every request is logged with tools, tokens, latency and feedback, and reviewed.
  • The full test suite runs on every change; every guard has a test naming its failure.

Adoption roadmap

  1. Weeks 1 to 4: informed. One system, read tools, delegated identity, rendering, interaction log. Measure what people ask.
  2. Weeks 5 to 8: analytical. Add the analytics engine for the questions managers ask most; add knowledge for the processes people ask about.
  3. Weeks 9 to 12: assisted. Add write tools with confirmation for two or three high-volume transactions.
  4. Next quarter: orchestrated. Workflow awareness, document intake, a second system.
  5. Later: autonomous. Specialist agents with standing goals, once evaluation and budgets are mature.

Enterprise AI agent: frequently asked questions

What is the difference between an enterprise AI agent and a chatbot?

A chatbot answers from text. An enterprise agent performs work in business systems on behalf of an identified person, within that person’s permissions, and grounds its answers in live records and computed results.

Does an enterprise agent need its own permissions?

No. It should hold no role of its own and act only as the person using it, so the business system’s permissions and audit trail apply unchanged.

Should business rules be written into the prompt?

No. Rules belong in the system of record, where they are enforced for everyone. The agent reads the outcome and reports any refusal exactly.

How does the agent produce accurate numbers and forecasts?

The language model chooses and explains the analysis; a local analytics engine fetches the complete data as the person and computes every number, forecast and anomaly in code. Only a summary returns to the model.

What is RAG used for in an enterprise agent?

For meaning: how processes work, what documents are for, what normally happens next, and the organisation’s own policies. Business data is always read live through tools.

Do I need LangChain or another agent framework?

Not necessarily. Frameworks help prototypes; production agents on systems of record benefit from direct control over prompts, tool selection, permissions and guards. Keep those parts in your own code either way.

Can one agent work across ERP, CRM and other systems?

Yes. With a canonical model and one connector per system, a single runtime can combine tools from several systems in one request, each call running with that system’s identity and permissions.

Is there an open-source implementation?

Yes. The Noviz agent is available on GitHub under AGPL-3.0 as a self-hosted reference implementation that uses ERPNext as its system of record.

Glossary

System of recordThe authoritative business system for a domain: ERP, CRM, HR, service desk.
Canonical entityA system-independent name for a business object, mapped by each connector.
ConnectorThe component that translates canonical operations into one system’s API.
ToolA typed operation the model may call, with a description and argument schema.
GatewayThe single function through which every tool call is authorized and executed.
Call specificationA generic operation the runtime asks an embedded add-in to run inside the system.
SpineThe small set of tools offered on every request, including tool discovery.
GuardA tested check between a proposed call and its execution that stops non-converging loops.
Data contractThe fixed columns each analytics dataset uses, independent of the source system.
Analytics bridgeThe typed tools and contract connecting LLM analysis to the local analytics engine.

Next steps

Questions or corrections: support@noviz.in