How to build an enterprise AI agent
This guide describes how to design, build and operate an AI agent that works inside the systems an enterprise already runs. It is written from the experience of building the Noviz platform, which serves agents for ERPNext and Microsoft Dynamics 365 Business Central from one common engine, and it explains each architectural decision together with the alternatives we considered and what production use taught us. The examples come from our open-source reference implementation, which uses ERPNext as its system of record; the architecture itself applies to any business system with an API.
In this guide
- A reference architecture for enterprise agents, with deployment topologies and the path of a single request.
- Sixteen building blocks, each with its purpose, the design we chose, the trade-offs and working code.
- How to run an agent in production: observability, security, cost, scaling and change management.
- The decisions that shape the platform, including why we built our own engine rather than using an agent framework.
- How to start from the open-source agent and extend it to your own systems.
Overview
Enterprises run on systems of record: the ERP that holds orders, stock and ledgers; the CRM that holds customers and pipelines; HR, payroll, procurement, service desks and document stores. Each of these systems already contains the organisation’s data, its business rules, its permissions and its approval processes. An enterprise agent is valuable when it can work with those systems the way an experienced employee does: find the right information, understand what it means, prepare and complete transactions, follow the approval process and explain the outcome.
The difficulty is not making a language model talk about business data. It is making an agent that is correct, that never exceeds the permissions of the person using it, that keeps cost predictable at scale, and that remains maintainable as the number of systems, customers and capabilities grows. Those are architecture problems, and they are the subject of this guide.
| Part | Contents |
|---|---|
| 1. Foundations | What an enterprise agent is, where it creates value, maturity levels, and the design goals that drive every later decision. |
| 2. Reference architecture | Layers, deployment topologies, the lifecycle of a request, and agents that span several systems. |
| 3. Building blocks | Channels, connectors, tools, identity and permissions, tool selection, prompt architecture, context, knowledge, analytics, workflow, documents, rendering, the reasoning loop, models and multi-tenancy. |
| 4. Operating in production | Observability and evaluation, security, cost, scaling and change management. |
| 5. Architecture decisions | Frameworks, multi-agent designs and technology choices. |
| 6. Build it yourself | The open-source reference implementation, how to extend it, a production checklist and an adoption roadmap. |
Part 1. Enterprise AI agent foundations
What an enterprise AI agent is
An enterprise agent is software that receives a goal in natural language, decides which operations in one or more business systems will achieve it, performs those operations on behalf of a specific person, and reports the result. Three properties distinguish it from a chatbot or a search assistant:
- It acts. It creates, updates, submits and approves records, sends messages and produces documents, not only answers.
- It is delegated. Every action is performed as an identified person, within that person’s permissions, and is recorded in the system’s own audit trail.
- It is grounded. Its answers come from live records and computed results, not from what the model remembers or estimates.
| Capability | Chatbot or search assistant | Enterprise agent |
|---|---|---|
| Source of answers | Model knowledge, indexed documents | Live records, computed analyses, indexed process knowledge |
| Actions | None | Transactions, approvals, messages, documents |
| Identity | Anonymous or a shared service account | The signed-in person, as the business system sees them |
| Numbers | Generated by the model | Computed by code, explained by the model |
| Rules and approvals | Described in a prompt | Enforced by the system of record |
| Accountability | A chat transcript | The system’s audit trail plus a structured interaction log |
Where enterprise agents create value
The same architecture serves every business function. What changes is the set of connected systems and tools.
| Function | Typical systems | Examples of agent work |
|---|---|---|
| Sales and customer | CRM, ERP selling | Qualify leads, prepare quotations, convert orders, chase overdue invoices |
| Procurement | ERP buying, supplier portals | Raise requests, compare supplier offers, track deliveries, match invoices |
| Finance | General ledger, banking, tax | Explain balances, prepare entries, analyse profit and cash, forecast results |
| Supply chain | Inventory, warehouse, manufacturing | Stock positions, reorder needs, slow stock, production progress |
| People | HR, payroll, attendance | Leave and expense requests, headcount and cost analysis |
| Service | Service desk, field service | Triage tickets, summarise history, update status, notify customers |
| Management | All of the above | KPI dashboards, period comparisons, exception reports, executive summaries |
Maturity levels
Organisations rarely move straight to autonomous agents. Each level builds on the guarantees of the previous one, and the architecture in this guide supports all four.
| Level | What the agent does | What must already be in place |
|---|---|---|
| 1. Informed | Answers questions from live data and process knowledge | Connectors, read tools, delegated identity, rendering |
| 2. Assisted | Prepares transactions and documents for a person to confirm | Write tools, confirmation, audit, refusal handling |
| 3. Orchestrated | Walks multi-step processes and approval workflows end to end | Workflow awareness, document chains, notifications |
| 4. Autonomous | Pursues standing goals on a schedule (collections, reorder, period close) | Service identities, budgets, exception routing, strong evaluation |
Design goals and trade-offs
Every decision in this guide serves one of six goals. They conflict in places, and the architecture is largely the record of how we resolved those conflicts.
| Goal | What it means | How the architecture achieves it |
|---|---|---|
| Correctness | Every fact and number is true | Live reads through tools; computation in code; system refusals passed through unchanged |
| Governance | The agent never exceeds the person’s authority | Delegated identity; layered authorization; confirmation before any commit |
| System authority | Business rules live in one place | Rules, validations and approvals stay in the system of record; the agent reads their outcome |
| Predictable cost | Cost per request is known and bounded | Per-request tool selection, layered prompts, caching, budgets |
| Portability | New systems without rewriting the agent | A canonical model and a connector contract; an engine that never imports system code |
| Evolvability | Improvements reach every customer safely | One shared engine, tests for every guard, a review loop over real traffic |
The most consequential of these is system authority. Early in the project it was tempting to encode business rules in prompts: discount limits, approval thresholds, mandatory fields. Rules expressed in a prompt are suggestions to a probabilistic model; they drift from the real configuration and they differ for every customer. We moved all of them back into the systems of record and designed the agent to read outcomes instead. The agent proposes; the system decides; the agent reports exactly what the system said.
Part 2. Enterprise AI agent reference architecture
Layers
The architecture has five layers. Requests flow down from the channel to the system of record; results flow back up. Each layer has a single responsibility, which keeps the parts that change often (prompts, tools, connectors) away from the parts that must stay stable (identity, authorization, execution).
| Layer | Responsibility | Changes how often |
|---|---|---|
| Channels | Collect the request, display structured results. No business logic. | Rarely |
| Agent runtime | Identity, tool selection, prompt and context, reasoning, authorization, rendering | Often (prompts, guards), behind tests |
| Capability services | Models, retrieval, exact analytics | Independently, behind stable contracts |
| Connectors | Translate canonical operations into each system’s API | When a system or mapping changes |
| Systems of record | Data, rules, permissions, workflows | Owned by the customer, never by the agent |
Deployment topologies
The same runtime supports four ways of being deployed. The choice depends on who operates the agent and where data may travel.
| Topology | How it works | Best for |
|---|---|---|
| Embedded with a hosted relay | A small add-in inside the business system runs the calls; the hosted runtime reasons and returns call specifications. Credentials never leave the system. | SaaS delivery to many organisations with no infrastructure on their side |
| Self-hosted single organisation | The runtime runs in the organisation’s own environment and calls the system’s API directly. | Full control, regulated environments, internal platforms |
| Hosted multi-system | One runtime with several connectors serves an organisation across ERP, CRM and service systems. | Cross-functional work, shared analytics |
| Hybrid | Reasoning hosted; execution, knowledge or analytics kept inside the organisation’s network. | Data residency requirements |
Noviz runs the first topology for ERPNext and Business Central, and publishes the second as open source. Both use the same engine; only the edges differ.
Lifecycle of a request
Consider the request “Which customer invoices are more than 30 days overdue, and what should we do first?”
- Identify. The channel sends the request with the organisation’s key and the person’s identity and roles as the business system reports them. The runtime resolves the organisation, its subscription tier and budget, and the tools this person may use.
- Resume or start. If the request continues a conversation, its saved turn state is loaded and the history is trimmed to a fixed budget. A new conversation may be answered from cache.
- Select tools. A deterministic router maps the request to business modules (here, receivables). Only the relevant, permitted tools are offered, plus a small always-available spine that includes a tool-search capability.
- Assemble the prompt. A stable core, the sections for the matched modules, the rule blocks of the offered tools, the organisation’s own policy, and a short session block.
- Add context. Conversation state and the closest knowledge passages, within budgets.
- Reason. The model returns either an answer or tool calls.
- Authorize. Each call passes the permission gateway. Reads proceed; writes become proposals for the person to confirm.
- Execute. The connector performs the call in the business system as the person. The system applies its own permissions and validations.
- Loop. Results return to the model, which decides the next step. Guards stop loops that cannot converge.
- Render and record. The final response is a structured object rendered identically in every channel, and the turn is written to the interaction log with tools, tokens, latency and later feedback.
Agents that span several systems
Most real work crosses system boundaries: a customer complaint in the service desk relates to an order in the ERP and a contact in the CRM. Because every connector exposes the same canonical operations, one runtime can hold several connectors and the model can combine their tools in a single turn. Three rules keep this safe:
- Identity per system. The person is resolved separately in each system, and each call runs with that system’s identity and permissions.
- Canonical names are unique. Entities are namespaced by system where they would collide (a CRM account and an ERP customer are different entities with an explicit link).
- No hidden joins. Cross-system matching is done by explicit tools with explicit keys, so a result can always be traced back to the records it came from.
Part 3. Building blocks of an enterprise AI agent
Each building block below follows the same structure: its purpose, the design we use, the decisions behind it, and code from the open-source reference implementation.
Channels and embedding
People should not leave their system to use the agent. The most effective channel is a chat surface embedded in the business application itself, opened from the same menu as the rest of their work, with no second sign-in. In the embedded topology, the add-in inside the business system has a second, more important job: it executes the agent’s data calls in the person’s own session.
A turn is an exchange of structured messages between the add-in and the runtime:
- The add-in sends the message, the person’s identity and roles, and the turn identifier, authenticated with the organisation’s key.
- The runtime reasons and, instead of calling the system itself, returns call specifications: generic operations such as list, get, insert, submit, run report.
- The add-in executes each specification through the system’s own API, as the person, and posts the results back.
- The runtime resumes from saved state and continues until the final response.
// Relay to plugin: "run these calls for me, as this user"
{
"turn_id": "6f1c0b2e-...",
"status": "calls",
"calls": [
{ "id": "c1", "method": "get_list", "doctype": "Sales Invoice",
"fields": ["name", "customer", "due_date", "outstanding_amount"],
"filters": [["docstatus", "=", 1], ["outstanding_amount", ">", 0],
["due_date", "<", "2026-09-08"]],
"order_by": "due_date asc", "limit": 50 }
]
}
// Plugin to relay: the results, exactly as ERPNext returned them
{ "turn_id": "6f1c0b2e-...", "results": [ { "id": "c1", "ok": true, "rows": [ ... ] } ] }
Why this design. We evaluated giving the runtime a service account with broad access. It is simpler, but it turns the runtime into a high-value target, bypasses the system’s own permissions and makes every record look as if one service user created it. Executing inside the person’s session gives native permissions and a native audit trail for free, and the organisation opens no inbound ports.
Design note
Keep the add-in thin: execute calls, display results, store settings. Every piece of reasoning placed in the add-in has to be released through each platform’s marketplace; every piece placed in the runtime reaches all organisations on the next deployment.
Connector layer and canonical model
The runtime never uses a system’s native names. It works with canonical entities such as
customer, sales_order or leave_application, and canonical fields such as
date, total and status. A connector translates them into the system’s own
objects and fields, and implements one interface that every part of the runtime depends on.
// Every connector implements one interface (excerpt from the open-source agent)
export interface SystemConnector {
loginWithPassword(identifier: string, password: string): Promise<UserCredential>;
getUserRoles(identifier: string): Promise<string[]>;
list(entityKey: string, credential: UserCredential,
params?: { filters?: Record<string, any>; limit?: number; sortBy?: string }): Promise<any[]>;
get(entityKey: string, credential: UserCredential, id: string): Promise<any>;
create(entityKey: string, credential: UserCredential, data: Record<string, any>): Promise<any>;
update(entityKey: string, credential: UserCredential, id: string, data: Record<string, any>): Promise<any>;
submit(entityKey: string, credential: UserCredential, id: string): Promise<any>;
aggregate(entityKey: string, credential: UserCredential, params: AggregateParams): Promise<any>;
runReport(reportKey: string, credential: UserCredential, filters?: Record<string, any>): Promise<any[]>;
}
// The ERPNext mapping: canonical entity -> native doctype and fields
sales_order: {
doctype: "Sales Order",
fieldMap: { id: "name", customer: "customer", status: "status",
total: "grand_total", date: "transaction_date" },
},
sales_invoice: {
doctype: "Sales Invoice",
fieldMap: { id: "name", customer: "customer", status: "status", total: "grand_total",
outstanding_amount: "outstanding_amount", due_date: "due_date", date: "posting_date" },
},
Source: backend/src/core/types.ts, backend/src/erpnext/entityMaps/selling.ts in the open-source agent
The contract has four parts:
| Part | Purpose |
|---|---|
| Operations | list, get, create, update, submit, aggregate, count, run report, document output, workflow description |
| Entity mapping | Canonical entity to native object; canonical field to native field; child tables for line items |
| Report mapping | Canonical report keys and filters to the system’s own reports |
| Capability hints | Which tools a given system or add-in can run, and any per-system shaping of a call |
A product application registers its connector once at start-up, before the engine loads, and the engine reads every system through a registry. When we brought the same engine to Microsoft Dynamics 365 Business Central, the work was a new connector, its entity maps and an add-in; the reasoning, prompts, tools, gateway and renderers were unchanged.
Pitfall. Canonical names are a product decision, not a technical one. Choose them from the business vocabulary your users speak, keep them stable, and make the connector, not the prompt, responsible for every system-specific detail.
Tool design
The model can act only through tools. A tool has a unique name, a description written for the model, a JSON schema for its arguments and a handler. Tools are grouped into modules and registered centrally at start-up.
export interface ToolDefinition {
name: string; // "<module>.<action>", globally unique
description: string; // written for the model: when to use it, what it returns
module: string;
parameters?: object; // JSON schema for the arguments
handler: (args: any, session: Session) => Promise<any>;
promptRules?: string[]; // guidance injected only when this tool is offered
}
export const knowledgeModule: MCPModule = {
name: "knowledge",
description: "ERP process knowledge and live approval workflow",
tools: [ knowledgeSearchTool, workflowDescribeTool ],
};
Source: backend/src/core/types.ts, backend/src/modules/knowledge/index.ts in the open-source agent
Most tools are generated rather than written:
| Factory | Input | Generated tools |
|---|---|---|
| Entity factory | An entity configuration: canonical fields and allowed operations | <entity>.list, .get, .create, .update, .submit |
| Report factory | A report configuration mapped to the system’s report | One tool per named report |
| Workflow factory | A state transition | One gated tool per transition |
Hand-written modules cover what configuration cannot: analytics, charts, reports and documents, schema discovery, inbox actions, document intake and knowledge retrieval.
Granularity. We tried both extremes. One generic “query anything” tool gives the model too much freedom and produces malformed queries; hundreds of narrow tools overwhelm it. The balance that works is typed entity tools for everyday operations, a small number of powerful typed tools (aggregate, join, report) for analysis, and discovery for everything else.
Descriptions are part of the product. A description that states limits and alternatives (“returns at most 50 rows; use analytics.aggregate for totals”) prevents more errors than any prompt instruction, because the model reads it at the exact moment it chooses the tool.
Identity, permissions and governance
Authorization is layered. Each layer can only remove options; none can add access that another layer denied.
- Organisation. The key resolves one organisation, its active term, its tier (read-only or read and write) and its budget.
- Person. The signed-in person and their roles, as reported by the business system.
- Role allowance. Roles resolve to the tools that person may use. A sales user is never offered payroll tools.
- Tool selection. Only relevant permitted tools are offered for this request; discovery never reaches beyond the grant.
- Gateway. Every call, from the model or from any other route, passes one function. Unknown or ungranted tools are rejected; writes become proposals.
- System of record. The call runs as the person, so the system applies its own permissions, validations and approval workflow.
// The single enforcement and execution point for every tool call
export async function callTool(session: Session, toolName: string, args: any) {
const allowed = session.allowed_tools.includes("*") || session.allowed_tools.includes(toolName);
if (!allowed) throw new ToolNotAllowedError(toolName);
const tool = moduleRegistry.findTool(toolName);
if (!tool) throw new Error(`Unknown tool: ${toolName}`);
return tool.handler(args, session); // the handler calls the connector as session.credential
}
Source: backend/src/core/gateway.ts (business-rule hooks omitted) in the open-source agent
A single entry point makes the permission model auditable: to verify it, you read one function. Confirmation is part of the same design. Any tool that creates, updates, submits, cancels or sends returns the exact document and values it intends to write; the person confirms in the interface; only then is the call executed.
Important
The agent must never hold a role of its own. If it can do something the person cannot, the permission model has a hole that no prompt can close.
Tool selection at scale
An agent for a full business system has several hundred tools. Offering all of them on every request is slow, expensive and less accurate, and model providers cap the number of tools per request. We select tools per request in three stages: role allowance, a deterministic module router, and a spine with discovery.
const RELAY_SPINE_TOOLS = new Set([
"tools.search",
"data_table.list",
"data_table.search_schema",
"database_engine.execute_query",
"report.find",
"analytics.catalog",
]);
export function selectRelayTools(tools: ToolDefinition[], prompt: string): ToolDefinition[] {
const modules = relayModulesFor(prompt);
const wantsCharts = detectExplicitModules(prompt).has("analytics");
return tools.filter((t) => {
if (RELAY_SPINE_TOOLS.has(t.name) || t.module === "context") return true;
if (wantsCharts && (t.name === "analytics.aggregate" || t.name === "chart.build")) return true;
if (modules.size > 0 && (t.name.endsWith(".list") || t.name.endsWith(".get")) && toolMatchesBusinessModule(t, modules)) {
return true;
}
return false;
});
}
Source: backend/src/core/toolRelevanceFilter.ts in the open-source agent
The router matches the request against keyword lists per business module and widens related pairs (a customer question
about invoices also loads selling, because that is where invoices live). The same router decides which prompt sections
are loaded, so the tools offered and the guidance given can never disagree. Anything not offered remains reachable through
tools.search; tools found that way stay available for the rest of the turn.
Why deterministic. We also built model-based tool selection, an extra model call that picks tools. It is useful for very broad roles such as system administrators and is kept for them. For department roles, the keyword router is faster, free, predictable and easy to test.
Prompt architecture
A single large system prompt is the most common source of cost and inconsistency in agent projects: every rule is paid for on every request, and rules for one domain interfere with another. We assemble the prompt from layers.
| Layer | Content | When included |
|---|---|---|
| Core | Identity, operating principles, tool guidance, display conventions | Always; identical bytes on every request |
| Write layer | How to propose changes and wait for confirmation | Only when the person may write |
| Module sections | Domain guidance per business module | Only for modules the router matched |
| Tool rule blocks | Guidance a specific tool needs | When that tool is offered, once per conversation |
| Organisation policy | The customer’s own instructions | New conversations, that organisation only |
| Session block | Date, company, person, roles | Always; always last |
export function buildSystemPrompt(prompt: string, canWrite: boolean, identity: AppIdentity, frappeUser: string, frappeRoles: string[]): string {
const realModules = [...relayModulesFor(prompt)].filter((m) => BUSINESS_MODULE_KEYS.includes(m));
const sections: string[] = [];
if (realModules.length === 0) {
sections.push(CORE_SYSTEM_PROMPT);
} else {
sections.push(THIN_CORE, AUTOFLOW_RULES); // AUTOFLOW: any real business-data turn wants the document-chain facts
}
if (canWrite) sections.push(WRITE_OPERATIONS);
const seen = new Set<string>();
for (const m of realModules) {
for (const section of MODULE_PROMPT_SECTIONS[m] || []) {
if (!seen.has(section)) {
seen.add(section);
sections.push(section);
}
}
}
sections.push(buildUserContext(identity, frappeUser, frappeRoles));
return sections.filter(Boolean).join("\n\n");
}
Source: backend/src/systemPrompt/index.ts in the open-source agent
The order is fixed for a reason. Model providers cache the beginning of a prompt that is byte-for-byte identical across requests and charge less for it. Keeping the stable core first and the per-person session block last lets most of every request hit that cache.
Where guidance goes
The open-source agent ships every prompt block empty, so each organisation writes its own. The examples below show the shape only.
| Guidance | Location | Example |
|---|---|---|
| Who the agent is | systemPrompt/core/identity.ts | "You are the operations agent for Acme." |
| How to choose tools | systemPrompt/core/toolDiscovery.ts | "For totals, use analytics.aggregate." |
| How to present results | systemPrompt/core/display.ts | "Show lists as tables." |
| Before any change | systemPrompt/core/writeOperations.ts | "Show the values and wait for confirmation." |
| One business area | systemPrompt/modules/<module>.ts | "A quotation becomes a sales order." |
| When a module loads | MODULE_KEYWORDS in systemPrompt/modules/index.ts | selling: ["quotation", "sales order"] |
| One tool | The tool’s promptRules | promptRules: [REPORT_RULES] |
| Per-message hints | systemPrompt/core/hints.ts | "Use a relative date filter." |
| Organisation policies | Administration console (no code) | A policy document |
// systemPrompt/modules/selling.ts
export const SELLING_MODULE = "A quotation becomes a sales order.";
// systemPrompt/modules/index.ts
selling: ["quotation", "sales order"],
Placeholder wording. A request that mentions “quotation” loads the selling section and the selling tools together.
export const REPORT_RULES = "Use reports for files, not chat tables.";
{ name: "report.generate", /* ... */ promptRules: [REPORT_RULES] }
Placeholder wording. The rule is sent once per conversation, the first time the tool appears.
Important
Treat prompts as production code: version them, review them, and run the full regression suite on every change. A rule added to fix one conversation can quietly break ten others.
Context and memory
Context is everything the model sees besides the request and the tools. We organise it in three tiers by cost.
| Tier | Source | Use |
|---|---|---|
| Hot | The current conversation, recently referenced records, the last result set | Always included, within a budget |
| Warm | Knowledge and policy passages closest to the request | Added automatically, within a budget |
| Cold | Everything else in the systems and the knowledge store | Fetched only through tools, when the model asks |
- Budgets per source stop one large source from crowding out the others.
- Bounded history. Conversations keep a fixed number of earlier turns. We drop older turns rather than ask the model to summarise them, because silent loss of detail in a summary is harder to detect than a clear limit.
- Large results stay on the server. The model receives a page and a summary; the full set remains available for paging, export and analysis.
- Memory is scoped to the session, not the person, so two sign-ins never inherit each other’s state.
Knowledge and retrieval (RAG)
Tools give the agent access to data. They do not tell it what the data means: that a delivery reduces stock, that a submitted invoice can no longer be edited, that a purchase order in “Pending Approval” is waiting for a specific role. Retrieval-augmented generation supplies that meaning from documents at the moment it is needed, without carrying it in every prompt.
| Source | Example content | Kept current by |
|---|---|---|
| Process documentation | Order to cash, procure to pay, stock movements, period close, request and approval flows, document lifecycle | Versioned files, re-indexed on change |
| Organisation documents | Policies, standard operating procedures, contracts | Re-indexed on every edit; can be deactivated |
| Live workflow definitions | The organisation’s own approval workflows: states, transitions, roles | Read from the system at request time, never copied |
Ingestion
- Parse each document and its metadata: title, module, the document types it describes.
- Chunk by meaning at section headings, splitting long sections with a small overlap so a sentence on a boundary keeps its context.
- Label each chunk “document title > section” so it remains understandable on its own.
- Embed and store in PostgreSQL with pgvector, linked to the source document; re-indexing replaces only that document’s chunks.
- Skip unchanged sources by checksum, so embeddings are paid for once.
/** Splits a document into labelled retrieval chunks: one per "## " section,
* long sections further split by the shared policy chunker. */
export function chunkKnowledge(doc: KnowledgeFile): KnowledgeChunk[] {
const sections: { name: string; text: string }[] = [];
let current = { name: "Overview", text: "" };
for (const line of doc.body.split("\n")) {
const h2 = line.match(/^##\s+(.+)$/);
if (h2) {
if (current.text.trim()) sections.push(current);
current = { name: h2[1].trim(), text: "" };
} else if (!/^#\s+/.test(line)) {
current.text += line + "\n";
}
}
if (current.text.trim()) sections.push(current);
const chunks: KnowledgeChunk[] = [];
for (const s of sections) {
const parts = chunkText(s.text.trim());
parts.forEach((part, i) => {
const label = `knowledge:${doc.slug}#${s.name}${parts.length > 1 ? ` (${i + 1}/${parts.length})` : ""}`;
chunks.push({ label, content: `${doc.title} > ${s.name}\n\n${part}` });
});
}
return chunks;
}
Source: backend/src/core/knowledgeStore.ts in the open-source agent
Retrieval
Knowledge reaches the model in two ways: automatically, as the closest passages in the warm context tier, and on demand through a knowledge.search tool filtered by module, whose results carry their source titles for citation.
select d.title, e.label, e.content, 1 - (e.embedding <=> $1) as score
from context_embeddings e
join knowledge_documents d on d.id = e.knowledge_document_id
where d.active and ($2::text is null or d.module = $2 or d.module is null)
order by e.embedding <=> $1
limit $3;
Source: backend/src/core/knowledgeStore.ts (search) in the open-source agent
Workflow awareness
Static documents describe the standard process; the organisation’s real approval process lives in the system and changes over time. A second tool, workflow.describe, reads it live: the configured states and transitions and, for one document, its current state and the actions the system allows this person now.
{
"doctype": "Purchase Order",
"has_workflow": true,
"states": [ { "state": "Pending Approval", "doc_status": "Draft" },
{ "state": "Approved", "doc_status": "Submitted" } ],
"transitions": [ { "from": "Pending Approval", "action": "Approve", "to": "Approved",
"allowed_role": "Purchase Manager", "condition": null, "self_approval": false } ],
"document": { "name": "PUR-ORD-2026-00087", "current_state": "Pending Approval",
"docstatus": 0, "actions_for_you": [] }
}
Output of workflow.describe (backend/src/erpnext/workflow.ts). The workflow definition is metadata; the document and the actions available on it are read as the user, through ERPNext’s own get_transitions.
Benefit. With both, the agent answers “why can’t I submit this order?” correctly: the knowledge base explains approval workflows; the live workflow shows that this order needs a role the person does not hold. No rule inside the agent made that decision.
Important
Do not use retrieval for business data or business rules. Data read from vectors is stale and ignores permissions; rules belong in the system. Retrieval is for meaning.
The analytics engine: a bridge between LLM analysis and local computation
Analysis is where enterprise agents most often fail and where they can create the most value. Managers do not only ask for records; they ask how the business is performing, where money is being lost, what next quarter looks like and which customers, products or suppliers need attention. A language model alone cannot answer these questions reliably: it is poor at arithmetic over many rows, it cannot read thousands of records within a prompt budget, and anything it estimates cannot be reproduced. A database query alone cannot answer them either: it does not understand the question, does not know which analysis fits, and cannot explain the result.
We therefore built analytics as a dedicated engine with two halves and a bridge between them. The LLM side understands the question, chooses the analysis and explains the result. The local analytics engine, running on our own servers next to the agent runtime, fetches the complete data and computes every number, comparison, forecast and anomaly exactly. The bridge is a small set of typed tools and a fixed data contract, so the two halves exchange intent and results, never raw rows.
What the engine is for
The engine carries a catalogue of 100 standard analyses in 12 categories. Each one is a complete, tested computation, not a template for the model to fill in.
| Area | Questions it answers | Examples of analyses |
|---|---|---|
| Financial results | Are we profitable, where is margin made or lost, how will next periods close? | Profit and loss, gross and net profit, profit margin, budget against actual, revenue against target, with forecasts |
| Cash and working capital | How much cash will we have, who owes us, whom do we owe? | Cash position, inflow and outflow, working capital, receivables and payables ageing, payment history |
| Stock flow | What do we hold, what moves, what is dead, what must be reordered? | Stock balance and valuation, turnover, stock ageing, dead stock, slow and fast movers, reorder analysis, ABC classes |
| Customers and sales | Who is growing, who is slipping, where is revenue concentrated? | Sales by customer, item, territory and salesperson, growth trends, profitability by customer, purchase frequency |
| Suppliers and purchasing | Who delivers late, who is getting more expensive, where are we dependent? | Supplier delivery performance, price comparison and trends, spend concentration, open orders |
| Operations and people | Are we efficient, where is waste, what does the workforce cost? | Production efficiency and wastage, workstation use, project budget and progress, payroll cost, turnover |
| Executive | How healthy is the business overall, and what needs attention first? | Business KPI dashboard, overall business health, executive summary with watch-outs |
Every analysis can add a forecast for the next periods, flag anomalous periods and locate the point where a trend changed, and any analysis can be compared across two periods or two segments. This is what lets the agent address complex business issues: not just “sales fell”, but which customers drove the fall, when it started, whether it is unusual, and what the next quarter is likely to look like if nothing changes.
How a request flows through the engine
- Understand. The model reads the question and identifies the intent, period and scope.
- Choose. It calls the catalogue tool to find the matching analyses by name and purpose, and asks the person when the request is ambiguous.
- Request. It calls the build tool (or the period or segment comparison tool) with an analysis key and parameters only.
- Fetch. The runtime retrieves each dataset the analysis needs, for the complete period, through the connector and as the person. The system’s permissions decide what is visible.
- Normalise. Rows are mapped to the data contract: fixed columns per dataset (date, party, item, amount, quantity, account). Every system maps to the same contract.
- Compute. The engine runs the analysis in Python: aggregation, grouping, period arithmetic, ageing, ranking, concentration, ABC classification.
- Forecast and detect. A forecast blends a damped-trend (Holt) model with linear regression and reports an 80% range; anomalies are found with a robust z-score based on the median absolute deviation; a change point is located with a Welch t-test across candidate splits.
- Interpret. An insight layer adds KPIs with change against the prior period, concentration and trend findings, watch-outs and recommendations.
- Render. Charts are drawn by the engine, assembled into a dashboard for the chat and a PDF for sharing.
- Explain. Only the summary and KPIs return to the model, which explains them, answers follow-up questions and proposes the next analysis or action.
Why the work is divided this way
| Property | LLM alone | LLM + local analytics engine |
|---|---|---|
| Accuracy | Numbers estimated from a sample of rows | Every number computed in code over the complete data |
| Reproducibility | Different answer on a different run | Same data and parameters give the same result |
| Scale | Limited by the prompt size | Limited by the server, with chunked fetching |
| Cost | Thousands of rows of tokens per question | A few hundred tokens of summary |
| Permissions | Depends on what was pasted into the prompt | Data fetched as the person; nothing they cannot see |
| Explanation | Fluent but unverifiable | Fluent and grounded in computed results |
The engine runs as its own service because it has a different language, a different load profile and a different failure mode from the reasoning runtime: heavy numerical work should never slow down a conversation, and a failed analysis must never take a conversation down with it. It listens only on a private interface, holds no system credentials and never contacts a customer system; the runtime brings it the data.
In the open-source agent
The reference implementation ships the same pattern at a smaller scale: one analysis, a monthly trend with growth and a projection, run as a local Python process.
"""Monthly trend: the sample analysis of the Python analytics engine.
Input (stdin, JSON): {"rows": [{"date": "2026-01-14", "total": 1200.0}, ...]}
Output (stdout, JSON): KPIs, a chart specification and observations.
The rows use canonical field names (date, total), the data contract every
connector maps its native fields to, so the analysis works for any ERP.
The rows never reach the language model; only this output does.
"""
import json
import sys
import numpy as np
import pandas as pd
def analyse(rows):
df = pd.DataFrame(rows)
if df.empty or "date" not in df or "total" not in df:
return {"kpis": [], "chart": None, "observations": ["No records in the period."]}
df["date"] = pd.to_datetime(df["date"])
df["total"] = pd.to_numeric(df["total"], errors="coerce").fillna(0.0)
monthly = df.set_index("date")["total"].resample("MS").sum()
growth = monthly.pct_change().mul(100).round(1)
# Linear projection for the next month (least squares over the series).
if len(monthly) > 1:
slope, intercept = np.polyfit(np.arange(len(monthly)), monthly.values, 1)
else:
slope, intercept = 0.0, float(monthly.iloc[0])
projection = max(0.0, float(slope * len(monthly) + intercept))
best = monthly.idxmax()
observations = [
f"Highest month: {best:%B %Y} at {monthly.max():,.2f}.",
f"The trend is {'rising' if slope > 0 else 'falling' if slope < 0 else 'flat'} by about {abs(slope):,.2f} per month.",
]
if growth.notna().any():
observations.append(f"Latest month-on-month change: {growth.iloc[-1]:+.1f}%.")
return {
"kpis": [
{"label": "Total", "value": round(float(monthly.sum()), 2)},
{"label": "Average per month", "value": round(float(monthly.mean()), 2)},
{"label": "Projected next month", "value": round(projection, 2)},
],
"chart": {
"type": "line",
"labels": [d.strftime("%b %Y") for d in monthly.index],
"series": [{"name": "Total", "values": [round(float(v), 2) for v in monthly.values]}],
},
"observations": observations,
}
if __name__ == "__main__":
print(json.dumps(analyse(json.load(sys.stdin).get("rows", []))))
Source: backend/analytics/monthly_trend.py in the open-source agent
handler: async (args, session) => {
if (!TREND_ENTITIES.includes(args.entity)) throw new Error(`entity must be one of: ${TREND_ENTITIES.join(", ")}`);
const rows = await systemConnector.list(args.entity, session.credential, {
filters: { date: ["between", [args.from_date, args.to_date]], status: ["not in", ["Draft", "Cancelled"]] },
limit: ROW_CAP,
sortBy: "date",
sortDir: "asc",
});
const result = await runPythonAnalysis("monthly_trend", rows.map((r) => ({ date: r.date, total: r.total })));
return {
entity: args.entity,
period: { from: args.from_date, to: args.to_date },
records_analysed: rows.length,
...(rows.length >= ROW_CAP ? { note: `Only the first ${ROW_CAP} records were analysed; narrow the date range for a complete result.` } : {}),
...result,
};
},
Source: backend/src/modules/trends/index.ts in the open-source agent
import { spawn } from "child_process";
import path from "path";
/**
* Runs one Python analysis (backend/analytics/<name>.py) as a child
* process: rows in on stdin, JSON result out on stdout. Keeps the agent a
* single deployment; a hosted platform can run the same scripts as a
* separate Python service without changing them. The rows go to Python,
* never to the language model; the model only sees the returned summary.
*/
export const ANALYTICS_DIR = path.resolve(__dirname, "../../analytics");
export function runPythonAnalysis(
name: string,
rows: object[],
opts: { command?: string; args?: string[]; timeoutMs?: number } = {}
): Promise<any> {
if (!/^[a-z0-9_]+$/.test(name)) return Promise.reject(new Error(`Invalid analysis name "${name}"`));
const command = opts.command ?? (process.env.PYTHON_BIN || "python3");
const args = opts.args ?? [path.join(ANALYTICS_DIR, `${name}.py`)];
return new Promise((resolve, reject) => {
const child = spawn(command, args, { timeout: opts.timeoutMs ?? 30000 });
let out = "";
let err = "";
child.stdout.on("data", (d) => (out += d));
child.stderr.on("data", (d) => (err += d));
child.on("error", (e) => reject(new Error(`Could not start the analytics engine (${command}): ${e.message}`)));
child.on("close", (code) => {
if (code !== 0) return reject(new Error(`Analysis "${name}" failed: ${err.trim() || `exit code ${code}`}`));
try {
resolve(JSON.parse(out));
} catch {
reject(new Error(`Analysis "${name}" returned invalid JSON`));
}
});
child.stdin.end(JSON.stringify({ rows }));
});
}
Source: backend/src/core/pythonAnalysis.ts in the open-source agent
To add an analysis, write a script that reads rows in the data contract and prints KPIs, a chart specification and observations, then expose it with a tool like analytics.monthly_trend. The model never needs to know how it is computed, only what it is for.
Workflow and approvals
Business work is a chain of documents: a request becomes an order, an order becomes a delivery, a delivery becomes an invoice, an invoice is paid. Systems of record track that chain through links between documents and enforce approval workflows at specific steps. The agent’s job is to walk the chain faster, never to shorten it.
- Read the state of the current document and its workflow, live, as the person.
- Propose the next document, created from its source so the system keeps links and progress tracking.
- Confirm with the person, showing the exact values.
- Let the system validate and apply its approval workflow; pass through any refusal unchanged.
- Leave approvals to approvers. An approver can act in the system or through the agent if they hold the role; the agent never approves on someone’s behalf.
- Report and offer the next step, with a notification where the workflow expects one.
Document intake
Incoming documents (supplier invoices, customer orders, receipts) are a natural task for an agent:
- The person uploads a file or takes a photo in the channel.
- A vision-capable model extracts text and structure: party, dates, line items, totals.
- The agent identifies the transaction type and matches the party and items in the system.
- It prepares a draft and presents it as a proposal; nothing is written until the person confirms.
Rendering and experience
The model does not produce HTML and does not draw charts. It chooses a response type, and the runtime produces a structured object that every channel renders with the same components.
| Type | Used for | Produced by |
|---|---|---|
| Text | Explanations, confirmations, system refusals | The model |
| Table | Lists of records with paging and links into the system | Table renderer |
| Cards | One record, an email or a notification with actions | Cards renderer |
| Chart and dashboard | Trends, breakdowns, KPI views | Chart builder or the analytics engine |
| File | Reports, the system’s own print formats, exports | Report generator or the system |
| Proposal | A write awaiting confirmation | The gateway |
{
"type": "table",
"title": "Overdue customer invoices",
"columns": ["Invoice", "Customer", "Due date", "Outstanding"],
"rows": [["ACC-SINV-2026-00412", "Gulf Traders LLC", "2026-08-14", 18450.00]],
"links": { "Invoice": "/app/sales-invoice/{value}" },
"page": { "offset": 0, "size": 20, "total": 37 }
}
Benefit. One response format means one design on every surface, whether the agent is embedded in ERPNext, in Business Central or in a web app, and the model cannot alter a number on its way to the screen because tables and charts are drawn from tool results.
Reasoning loop, guards and workers
The reasoning loop is a plain loop: select the tools for this step, inject their rule blocks, call the model, check the proposed calls, execute them, append the results, repeat. Tools found through discovery stay available for the rest of the turn.
const toolsForThisStep = () => {
const selected = selectRelayTools(allowedTools, relevanceSearchText);
const selectedNames = new Set(selected.map((t) => t.name));
const forced = allowedTools.filter((t) => forcedToolNames.has(t.name) && !selectedNames.has(t.name));
return [...forced, ...selected].slice(0, MAX_TOOLS_PER_REQUEST);
};
for (let i = 0; i < appConfig.llm.maxToolIterations; i++) {
const tools = toolsForThisStep();
for (const t of tools) for (const block of t.promptRules ?? []) injectPromptBlock(messages, block);
const response = await this.llm.chat(messages, tools);
if (response.tool_calls.length === 0) { finalText = response.content || ""; break; }
messages.push({ role: "assistant", content: response.content || "", tool_calls: response.tool_calls });
for (const call of response.tool_calls) {
// guards: repeated calls, denied tools, step and time budgets ...
const result = await callTool(session, call.name, call.arguments);
if (call.name === "tools.search") {
for (const name of foundToolNames(result).slice(0, TOOLS_SEARCH_FORCE_CAP)) forcedToolNames.add(name);
}
messages.push({ role: "tool", content: JSON.stringify(result), tool_call_id: call.id, name: call.name });
}
}
Source: backend/src/core/reasoningEngine.ts (shortened) in the open-source agent
Guards. Models sometimes repeat a call that cannot succeed, such as paging through thousands of rows to build a total. Guards detect these patterns and redirect the model to the right tool, an aggregate or a report. Each guard exists because of a reproduced failure and has a test that names it; we remove a guard only when that failure is proven impossible.
State and workers. The runtime is stateless between steps: conversation and turn state live in the database and cache. Several workers serve requests behind one endpoint and any worker can continue any turn. Long analyses and reports can move to a background queue so that web requests return quickly.
Why plain functions. We kept the loop as ordinary, named functions rather than a graph or framework abstraction, so that each step can be read, tested and traced on its own, and so that a guard can sit exactly between “the model proposed” and “the tool executed”.
Model connectivity and strategy
The runtime talks to models through one provider interface. Any service implementing the chat-completions format with tool calling can be used.
export interface LLMProvider {
chat(request: {
messages: ChatMessage[];
tools?: ToolSchema[];
temperature?: number;
maxTokens?: number;
}): Promise<{ content: string | null; toolCalls: ToolCall[]; usage: TokenUsage }>;
}
Illustrative. The open-source implementation is backend/src/providers/llm/openaiProvider.ts
- Switchable at runtime. Base URL, model and key are settings; changing provider takes seconds and no deployment.
- One model per task. Reasoning, document vision and embeddings can use different providers, so one provider’s limit or outage does not stop everything.
- Timeouts, retries and clear failures on every call.
- Rate and budget gates. A shared token-rate gate keeps all workers under provider limits; a monthly budget per organisation bounds cost.
- Usage recorded per request: prompt, cached and completion tokens, which is how cost problems are found and fixed.
Multi-tenancy and customization
A hosted agent serves many organisations from one runtime. Four mechanisms keep them separate and individual:
| Mechanism | What it isolates or adapts |
|---|---|
| Organisation keys and tiers | Who is calling, whether they may write, their term and budget |
| Scoped data | Turn state, logs, usage, knowledge and caches are keyed by organisation |
| Organisation policy | Plain-language instructions that apply only to that organisation |
| Customization layers | Standard behaviour, then a system layer, then a client layer that can extend, override or add, and switch modules on or off; the base is never edited |
Benefit. One engine and one release serve every organisation, yet each can have its own behaviour without a fork, the same way ERPNext and Business Central themselves are customised.
Part 4. Running an enterprise AI agent in production
Observability and evaluation
An agent is only as reliable as the evidence about how it behaves. We rely on four instruments:
- Interaction log. Every request records the prompt, the modules matched, the tools offered and called, the gates that fired, the response type, latency, token usage and later user feedback. It is the single source for every improvement.
- Review loop. Negative feedback, blocked gates and slow or failed turns are reviewed across all organisations. A pattern found in one organisation’s traffic is fixed once, centrally, and every organisation receives the fix on the next release.
- Automated tests. The engine has more than 1,300 automated tests covering tool generation, permissions, prompt assembly, tool selection, guards, renderers and connector mapping. Every change runs the full suite.
- Regression discipline. Before a guard, prompt layer or session behaviour is changed, its history is read and the change is shown not to reopen the failure it was added for.
Security and compliance
| Concern | Control |
|---|---|
| Excess privilege | Delegated identity; layered authorization; no agent role |
| Unwanted changes | Confirmation before every commit; read-only tier |
| Credential exposure | Embedded execution: credentials never leave the system. Self-hosted: per-person credentials encrypted at rest |
| Prompt injection | Content from records, emails and documents is treated as data, never as instructions; it cannot grant tools, change permissions or skip confirmation, because those are enforced in code |
| Data leakage between customers | Every store and cache keyed by organisation; retrieval filtered by organisation |
| Data minimisation | Large result sets and analytics rows stay on the server; the model receives pages and summaries |
| Mailbox safety | Mail is read without deletion or moving |
| Exposure | Services listen privately behind one gateway; no inbound access to customer systems |
| Accountability | System audit trail as the real person; interaction log per request |
Cost management
Cost per request is an architectural property. The largest savings in our platform came from design, not from cheaper models:
- Tool selection sends a few dozen tool schemas instead of hundreds.
- Layered prompts send only the guidance a request needs, with a stable prefix that providers cache.
- The analytics bridge replaces thousands of rows with a summary of a few hundred tokens.
- Response caching answers repeated first questions exactly and very similar ones semantically, scoped per organisation.
- Budgets and gates make the worst case bounded and visible.
Scaling
- Modular monolith first. The runtime is one service with clear modules. A part becomes its own service only when it has a different language, load or failure profile: analytics and knowledge qualify; the reasoning hot path does not, because a network hop between deciding and acting adds latency and drift.
- Stateless workers behind one endpoint, with state in PostgreSQL and Redis.
- Background turns for long analyses and reports.
- Provider capacity monitored per minute, with headroom planned before limits are reached.
Change management
- One source of truth. What runs in production is exactly what is in the main branch.
- Tested deployments. Tests first, build beside the running version, swap, health check, automatic rollback, nothing left behind.
- Prompts and guards as code, reviewed and tested like any other change.
- Narrow changes. Fix the root cause in the smallest part that owns it; avoid patching symptoms in the prompt.
Part 5. Architecture decisions
Frameworks: why we built our own engine
LangChain, LangGraph, CrewAI, AutoGen and Semantic Kernel are valuable for prototypes and for applications that mainly connect many third-party sources. We evaluated them and chose a small engine of our own, for reasons specific to agents that act on systems of record.
| Concern | With a general framework | With our own engine |
|---|---|---|
| Prompt control | Prompts assembled by templates; exact bytes hard to see and stable across versions | Every byte and its order controlled, enabling prompt caching and reproducible behaviour |
| Token cost | All tools and long default instructions | Per-request tool selection and layered prompts |
| Permissions | Added around the framework | One gateway is the only path to a tool |
| Guards | Hard to place between proposal and execution | Named, tested functions in the loop |
| Debugging | Failures deep inside framework stacks | A short loop; each step visible in the log |
| Dependencies | Frequent breaking changes, large dependency tree | A few mature libraries that change rarely |
| Capacity | Frameworks organise code; they do not add capacity | Capacity from workers, caching and provider limits we manage directly |
If you start with a framework, keep the gateway, the connector contract and the tool definitions in your own code. Then orchestration can be replaced later without rebuilding the system.
Multi-agent designs
Several agents talking to each other multiply token use and latency, and make failures harder to trace. For conversational work in business systems, one agent with many well-selected tools performs better. Specialist agents earn their place when they have standing goals of their own, such as collections, replenishment or period close, or when work runs in parallel for a long time. We build them as instances of the same engine with their own tool set and prompt layer, behind the same gates, rather than as a separate agent framework.
Technology choices
| Area | Choice | Reason |
|---|---|---|
| Runtime | Node.js, TypeScript, Express | Typed tool schemas and connectors; strong concurrency for API-bound work |
| Data and vectors | PostgreSQL with pgvector | One database for state, logs, settings and vectors |
| Cache and coordination | Redis | Response cache, rate limits, token-rate gate across workers |
| Analytics | Python: pandas, NumPy, SciPy, matplotlib | The standard numerical stack, run as a private service |
| Documents | PDF generation and text extraction libraries | Reports out; uploaded PDFs and Word files in |
| Interfaces | Embedded pages in each system; React for web | One response format rendered natively |
| Tests | Jest; pytest for analytics | Fast tests for every component |
Part 6. Build your own enterprise AI agent
The open-source reference implementation
The Noviz agent is published as open source under AGPL-3.0 as a self-hosted reference implementation of this architecture. It uses ERPNext as its system of record and includes the connector contract and ERPNext connector, the tool factories and registry, the gateway and role-based permissions, per-request tool selection, the layered prompt pipeline (shipped empty for your own guidance), context tiers, knowledge retrieval with live workflow awareness, a Python analytics sample, renderers, reports, document intake and the agent and administration web apps.
github.com/Zeetech-soft-solution/erp-ai-agent
Prerequisites
- A system of record with an API (the reference uses ERPNext 15 or 16 and an API key for a service user)
- Node.js 20 or later, PostgreSQL 15 or later with the
vectorextension, Python 3 with pandas and NumPy - A key for an OpenAI-compatible model provider
Install
git clone https://github.com/Zeetech-soft-solution/erp-ai-agent.git
cd erp-ai-agent/backend
npm install
cp .env.example .env # ERPNext URL and key, LLM key, DATABASE_URL, CREDENTIAL_ENCRYPTION_KEY
# Create the schema: run every migration in filename order
for f in src/db/migrations/*.sql; do psql "$DATABASE_URL" -f "$f"; done
npm test # the full test suite
# Optional: knowledge base and Python analytics
npm run knowledge:index # embed backend/knowledge/**/*.md
pip install -r analytics/requirements.txt # pandas and numpy for analytics/*.py
npm run dev # API on the port set in .env
cd ../frontend/agent && npm install && cp .env.example .env && npm run dev
cd ../admin && npm install && cp .env.example .env && npm run dev
Code structure and why it is organised this way
backend/src/
core/ engine: gateway, reasoning loop, module registry, factories,
context assembly, tool filtering, renderers registry, stores
config/ canonical entities, role policy, reports, workflows, settings
erpnext/ the ERPNext connector: client, entity maps, report map
modules/ hand-written tools: analytics, trends, chart, reports, documents,
schema, tool discovery, inbox actions, context, knowledge
providers/ LLM, embeddings, vision, context tiers
renderers/ table, cards, chart
systemPrompt/ layered prompt: core blocks and module sections
routes/ auth, agent, admin, tools, policy documents, webhooks
db/migrations/ PostgreSQL schema
backend/knowledge/ process documents for the knowledge base (RAG)
backend/analytics/ Python analyses, run as a child process
frontend/
agent/ the user-facing agent app (React)
admin/ the administration console (React)
Each area of the code exists for a reason, and the boundaries between them are what make the agent portable and maintainable.
| Area | Purpose | Benefit of keeping it separate |
|---|---|---|
core/ | The engine: gateway, reasoning loop, registry, factories, context, tool selection, stores | Never imports a system’s code, so it works unchanged for any system of record |
core/gateway.ts | The single path to every tool | Permissions can be audited by reading one function |
config/ | Canonical entities, role policy, reports, workflows, settings | New coverage is configuration, not code, and is reviewed as data |
erpnext/ | The connector: client, entity maps, report map, workflow reader | Everything system-specific in one folder; a second system is a sibling folder |
modules/ | Hand-written tools: analytics, trends, charts, reports, documents, schema, discovery, inbox, knowledge | Each capability can be added or switched off on its own |
systemPrompt/ | Layered prompt: core blocks, module sections, hints, keywords | Guidance changes without touching logic; loaded only when relevant |
providers/ | Models, embeddings, vision, context tiers | A provider can be replaced by writing one class |
renderers/ | Table, cards, chart | One output format for every channel |
routes/ | HTTP surface: sign-in, agent, admin, tools, knowledge | Transport kept away from reasoning |
db/migrations/ | The database schema, in order | Reproducible environments |
knowledge/ | Process documents for retrieval | Meaning maintained as documents, not prompt text |
analytics/ | Python analyses run by the engine | Exact computation in the right language, outside the model |
frontend/ | Agent and administration apps | Interfaces that only render structured responses |
Extend the agent
| Task | What to change |
|---|---|
| Add an entity | Add its configuration (canonical fields and operations) and its mapping in the connector. The factory generates the tools. |
| Grant it to a role | Add the tool names to that role in the role policy. |
| Add a custom tool | Write a module with a definition and handler and register it; the gateway applies permissions automatically. |
| Connect another system | Implement the connector interface and entity maps for it (a CRM, HR or service desk); tools, prompts, gateway and interfaces stay unchanged. |
| Add process knowledge | Add a Markdown document with title, module and sections under knowledge/ and index it. |
| Add an analysis | Add a script under analytics/ that reads the data contract and returns KPIs, a chart and observations; expose it with a tool. |
| Write guidance | Fill the prompt blocks described in Where guidance goes. |
| Change model provider | Set base URL, model and key. |
Production checklist
- Every tool call passes one gateway; no handler is called directly.
- Every system call runs as the signed-in person; a service identity is used only for metadata.
- Writes require explicit confirmation; read-only access is available.
- Business rules live in the system of record; refusals are shown as returned.
- Numbers and analyses are computed by code; rows are not sent to the model.
- Tools are selected per request; prompts are layered with a stable prefix.
- Knowledge explains processes; live data and workflows are read through tools.
- Content from records, emails and files is treated as data, never as instructions.
- Timeouts, retries, rate limits and budgets are configured.
- Every request is logged with tools, tokens, latency and feedback, and reviewed.
- The full test suite runs on every change; every guard has a test naming its failure.
Adoption roadmap
- Weeks 1 to 4: informed. One system, read tools, delegated identity, rendering, interaction log. Measure what people ask.
- Weeks 5 to 8: analytical. Add the analytics engine for the questions managers ask most; add knowledge for the processes people ask about.
- Weeks 9 to 12: assisted. Add write tools with confirmation for two or three high-volume transactions.
- Next quarter: orchestrated. Workflow awareness, document intake, a second system.
- Later: autonomous. Specialist agents with standing goals, once evaluation and budgets are mature.
Enterprise AI agent: frequently asked questions
What is the difference between an enterprise AI agent and a chatbot?
A chatbot answers from text. An enterprise agent performs work in business systems on behalf of an identified person, within that person’s permissions, and grounds its answers in live records and computed results.
Does an enterprise agent need its own permissions?
No. It should hold no role of its own and act only as the person using it, so the business system’s permissions and audit trail apply unchanged.
Should business rules be written into the prompt?
No. Rules belong in the system of record, where they are enforced for everyone. The agent reads the outcome and reports any refusal exactly.
How does the agent produce accurate numbers and forecasts?
The language model chooses and explains the analysis; a local analytics engine fetches the complete data as the person and computes every number, forecast and anomaly in code. Only a summary returns to the model.
What is RAG used for in an enterprise agent?
For meaning: how processes work, what documents are for, what normally happens next, and the organisation’s own policies. Business data is always read live through tools.
Do I need LangChain or another agent framework?
Not necessarily. Frameworks help prototypes; production agents on systems of record benefit from direct control over prompts, tool selection, permissions and guards. Keep those parts in your own code either way.
Can one agent work across ERP, CRM and other systems?
Yes. With a canonical model and one connector per system, a single runtime can combine tools from several systems in one request, each call running with that system’s identity and permissions.
Is there an open-source implementation?
Yes. The Noviz agent is available on GitHub under AGPL-3.0 as a self-hosted reference implementation that uses ERPNext as its system of record.
Glossary
| System of record | The authoritative business system for a domain: ERP, CRM, HR, service desk. |
| Canonical entity | A system-independent name for a business object, mapped by each connector. |
| Connector | The component that translates canonical operations into one system’s API. |
| Tool | A typed operation the model may call, with a description and argument schema. |
| Gateway | The single function through which every tool call is authorized and executed. |
| Call specification | A generic operation the runtime asks an embedded add-in to run inside the system. |
| Spine | The small set of tools offered on every request, including tool discovery. |
| Guard | A tested check between a proposed call and its execution that stops non-converging loops. |
| Data contract | The fixed columns each analytics dataset uses, independent of the source system. |
| Analytics bridge | The typed tools and contract connecting LLM analysis to the local analytics engine. |