Your team spends its month matching lines, chasing breaks and re-keying the same data, then runs out of time for the work that actually needs a qualified person. Velocyt builds agents that absorb the volume without ever taking a decision your auditor would need to test.
Your ERP already matches the clean lines. What it cannot do is work the residual: the deposit that nets against forty transactions, the intercompany balance where both sides are internal and either could be wrong, the invoice with no purchase order and a coding call to make. That is where the hours go, and it is also where the risk concentrates, because a break that gets cleared incorrectly disappears into a set of books that still balance.
Most automation sold into finance targets the clean match, which was already solved, and leaves the residual alone. We do the opposite. The agent earns its place on the judgment-adjacent work, under controls tight enough that a wrong proposal cannot become a wrong entry.
An AI that proposes and then clears its own item has collapsed maker and checker into one actor, which is the first thing an auditor will find. So the split is structural: the model proposes an association or a classification, a deterministic rule recomputes whether that proposal actually reconciles, and only the rule clears. Every disposition reconstructs as four links, the inputs the model saw, its advisory output, the rule that acted, and the human sign-off where it crossed the gate.
The second half matters as much. Controls should be calibrated, not maximised. Payables is high volume and low consequence, so clean matched invoices under threshold run straight through and only the judgment routes to a person. A reserve estimate or a revenue determination is the opposite, and there the AI stays entirely outside the number. A vendor who applies the same control weight to both has not thought about either.
Anything above materiality goes to a person. So does any break the agent cannot explain with a check-confirmed cause, however small.
FX is recomputed from the period rate, never inferred. Duplicate detection runs on fixed keys. The arithmetic is never the model's opinion.
The auto-cleared population is risk-weighted and sampled, because the dangerous failure is a wrong match that sums correctly and tolerance cannot catch it.
The hardest version of matching, because both sides are internal and either can be wrong. The agent proposes which entity erred and why, the rule recomputes FX deterministically from the period rate so a remeasurement swing cannot hide a real break, and a person owns unmatched intercompany at period end. The failure mode worth naming out loud is an uneliminated balance leaking into consolidated figures, which is a completeness failure rather than a matching one, and it is checked as its own control.
Capture, three-way match, duplicate detection on multiple keys, and proposed general ledger coding with rationale for non-PO invoices. Clean matched invoices under threshold run straight through, because the consequence ceiling here is genuinely low. Payment authorization stays human and stays behind your existing approval matrix, always. Banking-detail changes are flagged by the agent and verified by a person out of band.
Recurring and templated entries stay deterministic, because they never needed a model. The agent earns a place on judgment entries only, assembling the scattered evidence to propose an accrual with its support, and reviewing manual entries for anomalies. It never posts. A person owns every posting that involves judgment.
The decomposition into price, volume, mix, rate and FX is fixed arithmetic, so the rules compute the split and the agent narrates it. The explanation must cite the computed movement rather than generate a plausible story, because looks-explained is not correctly-explained. Your reviewer checks the support, not just the prose.
Manual top-side entries are the classic misstatement vector, and they are hard to review at volume. The agent flags unusual account combinations, round numbers, post-close timing and unusual preparers. It flags only. It never blocks and it never posts, and known patterns stay on deterministic rules so the model is reserved for the unknown ones.
The open question in this field is not whether an agent can act. It is under what evidence it is defensible to let it. We are building the answer in public as a pre-registered evaluation harness that scores an agent's action decisions against PCAOB AS 2201 severity tiers, with under-escalation weighted by the materiality band rather than by intuition.
Not autonomy. A published deployment contract, written before the agent is tested, that names the conditions under which the system may auto-clear one specific class of item: zero material-weakness events, the deterministic anchor checks passing, and consistent action behaviour across repeated trials with the confidence bound reported. Where the current design contemplates this at all, it is for items below the inconsequential threshold, on low-risk tiers.
A Tier-1 action evaluation is sufficient for routine monitoring and regression gating. It is not sufficient on its own to clear an agent for unbounded authority on materially significant controls. That needs expert red-teaming and human comparison studies, which are out of scope for the current version and named as such in the pre-registration rather than glossed over. If a vendor tells you their agent is cleared to act and cannot show you the standard it cleared, there is no standard.
Velocyt is Rizwan Ahmed, ACCA. Audit and enterprise finance before the engineering, which is where the posture comes from: a conclusion is only as good as the evidence trail behind it. The claims on this page are not case studies you cannot check. They are open-source repositories with pinned model versions, checked-in evaluation results and hand-graded outputs.
The retrieval layer was built and measured on US consumer-debt law, where a fabricated citation is both an engineering failure and a regulatory one. Every cited chunk is structurally bound to a retrieved chunk and validated after generation by a deterministic set-membership check, and below a tuned confidence threshold the system refuses with a typed reason rather than answering. That property is enforced in code, not requested in a prompt.
Alongside it: a governance platform separating deterministic policy enforcement from advisory model output, with segregation of duties enforced at the database layer rather than in application code, and an independent benchmark of AI-writing rulesets built on a seeded ground truth across a 111,000-word corpus built from eighteen SEC filings, each with its accession number recorded. Both public. See the repositories →
Forecasting. It is judgment-heavy and the right answer is unknowable, which makes it the worst possible candidate for handing authority to a model. There is a defensible narrow slice, assembling inputs and drafting variance-to-forecast commentary, but the call stays with your forecaster and we would sequence it last.
A bookkeeping agent. We will refuse the monolithic framing. Granting a model authority over the ledger is the opposite of everything above. Decomposed, it is workflows you already have: coding is payables, entries are close, reconciliations are matching. Authority gets placed per workflow, never over the books. For a small business with no audit the calculus is different, and that is worth saying plainly rather than pretending the rule is universal.
Bank reconciliation, mostly. Amount, date and reference matching is solved in every ERP and in the tools you already pay for. A model earns a thin slice at the ambiguous break, fuzzy matching when references are garbled and classifying an unexplained difference. If someone is quoting you for a bank rec agent, ask what share already auto-matches.
Neither is something these systems do, by design. If that is what you are shopping for, we are not the right fit and we will say so on the first call.
A workflow audit is a short, specific review of where your close spends time a system could absorb, and which parts should stay exactly where they are. No obligation, no pitch deck. You leave with a map of your own process either way.
Accounting or advisory firm, and want this running under your own brand for your clients? Tell us.