Platform · Measure · The ledger

Analytics

Six months into most AI rollouts, nobody can tell the CFO what it costs per outcome or what it replaced — finance wants chargeback, the board wants proof, and the vendor's dashboard shows API calls. Here every run is metered natively — tokens and dollars per message, per call, per worker — and rolls up into scorecards you can defend a budget with: cost per successful run, spend against a hard cap, exportable by group.

See your fleet's numbers Budgets & governance
At a glance
  • Meteringtokens + USD, per message
  • Cost accountingper successful run
  • Depthfleet → entity → run
  • Entity screens6 types
  • Budgetsforecast vs hard cap
  • Any viewrole-scoped · exportable
IN A SOLUTION PACK

In a pack, the ledger arrives wired: runs, cost, and outcomes tracked from the first install.

See Solution Packs →
app.turtleaicoworker.com/statistics
The Statistics org overview: five KPI tiles with prior-period deltas, cost trend stacked by module, Top-by-spend leaderboard, activity heatmap and the attention feed
The org overview: five health KPIs, cost trend, spend leaderboard, activity heatmap, attention feed. One scroll, no clicks.
Why it's built this way

If you can't price the outcome,
you can't defend the budget.

AI spend without outcome accounting dies in the next budget cycle — not because it didn't work, but because nobody could prove it did. The pilot's champion is left arguing anecdotes against an invoice.

So the accounting is native, not bolted on. Every model call records its tokens and dollars per message on the way through; every run totals cost and duration; every tool call lands in the audit trail. Analytics is a read of that ledger — attributable down to the worker, the run and the message, and exportable in the shape finance asks for.

What you can measure

Eight lenses. One ledger.

From the metering underneath to the export finance takes away: every layer between a token and a defensible number.

01 · Metering

Costed at the message, so it's honest at the quarter.

Roll-up numbers are only as good as what's underneath, and most AI reporting is built on estimates. Here the accounting is native: every model call records its tokens and its dollars per message, every run totals tokens, cost and duration, and every tool call is classified and logged in the audit trail. The quarterly number is a sum of real rows, not a model of a model — which is why finance can lean on it.

  • Per-message accounting: tokens and USD recorded as the run happens
  • Per-run totals: tokens, cost, duration — joined to the audit trail
  • Nothing self-reported, nothing sampled: a sum, not an estimate
A single run's cost breakdown: per-message token counts and USD amounts rolling up to the run total
The atom of the ledger: one run, costed message by message
02 · Fleet health

The rollout's health, without a meeting.

Six months in, 'how is the AI going?' is usually answered with anecdotes. The org overview answers it in one scroll: five health KPIs with prior-period deltas, a spend trend, an activity heatmap, and a live attention feed of failures, tripped breakers, budget warnings and flagged PII. Whether spend is drifting, runs are failing or adoption is stalling is visible before anyone opens a ticket.

  • Five KPIs with Δ vs the prior period: success, spend, runs, adoption
  • A spend trend beside a top-by-spend leaderboard, so drift has a source
  • An attention feed: failures, tripped breakers, budget warnings, flagged PII
The org overview dashboard: KPI tiles with deltas, the spend trend, activity heatmap and attention feed in one scroll
One scroll, no clicks: the state of the rollout
03 · Spend

Cost drift gets a name before the invoice does.

The month the bill jumps, the vendor's answer is 'usage went up.' Here the spend trend sits beside a Top-by-spend leaderboard that ranks your most expensive workers. When the trend ticks up week-over-week, the worker driving it is one glance away — a specific agent, team or employee you can open, inspect and fix.

  • Daily spend with the workers behind it: agents, teams, employees
  • Top by spend ranks the most expensive workers, most costly first
  • From 'the bill went up' to the responsible worker in one click
shot: analytics-top-spendThe spend view: the spend trend beside the Top-by-spend leaderboard of named workers with their run counts and costs
The leaderboard that answers 'who spent it'
04 · Unit economics

Cost per outcome, not cost per attempt.

Cost per API call flatters everyone. A single worker's scorecard counts cost per successful run — total spend divided by the runs that actually finished the job — alongside a failure-reason breakdown, a p50/p95 duration trend and a trigger-source split. That distinction is where a prompt change that quietly doubled retries finally shows up: cost per attempt barely moves, cost per success jumps.

  • Cost per run and cost per successful run, side by side
  • Failure reasons named — a tool quota, a timeout — not guessed
  • p50/p95 duration trend and a trigger-source split
A single agent's scorecard: cost per run vs cost per successful run, failure-reason breakdown, p50/p95 duration trend
The honest number: what a finished job actually costs
05 · Compare

A leaderboard for every kind of worker.

Averages hide the worker that's quietly burning the budget. Entity Analytics ships one Statistics screen per type — Agents, Teams, AI Employees, Models, Users, Groups — each ranking its instances against each other. The Agents screen alone carries four leaderboards: most run, most expensive, highest failure rate, slowest. The interesting worker is the one whose rank differs across them: mid-pack on runs, first on failures.

  • Four agent leaderboards: most run · most expensive · highest failure · slowest
  • Six entity types, each ranked against its own kind
  • Bars turn warn/bad when a value crosses a threshold
The Agents statistics screen: four leaderboards ranking the same agents by runs, spend, failure rate and duration
The same agents, ranked four ways
06 · Teams & employees

Proof a team collaborates and an employee improves.

The Teams screen shows how each team actually performs — runs, cost and reliability per team, with member-level attribution. The AI Employees screen tracks whether a named employee is earning autonomy: a task funnel, a human-intervention rate that should be falling, and a memory panel that should be growing. A performance review becomes a chart, not an argument.

  • Per-team scorecards with member-level attribution
  • A task funnel: assigned, attempted, completed, escalated
  • A declining intervention rate as an employee earns autonomy
The Teams statistics screen: per-team scorecards with runs, cost and reliability for teams with real history
How the work moves, and whether the worker is learning
07 · Models

Migrate models on evidence, not a hunch.

Model choice is a recurring cost decision most teams make once and never revisit. The Models screen weighs every LLM you run by cost, latency and success rate, side by side, measured from your own traffic — not from a vendor benchmark. When a cheaper model would hold the line on an extraction step, the numbers say so, and the swap is a per-agent setting, not a rebuild.

  • A performance table: cost · latency · success, per model, from your runs
  • Evidence for a migration before you commit traffic to it
  • Catch the expensive model doing cheap work, and vice versa
The Models statistics screen: per-model table comparing cost, latency and success rate measured from the workspace's own runs
The migration case, made from your own traffic
08 · Budgets & export

Spend against a ceiling, and receipts that leave the app.

Numbers that stay in a dashboard don't survive a budget meeting. Spend here reads against the hard caps set in Governance — per workspace, employee or team, with a period-end forecast and the projected breach date in view. And any scoped view exports: a group's spend for chargeback, an agent's reliability for a post-mortem, a personal scorecard for a review. The date window, granularity and scope carry into the export.

  • Spend vs hard caps, with period-end forecast and breach date
  • Views are role-scoped: org-wide for Owners/Admins, own assets for Builders
  • Any scoped view exports — chargeback, post-mortem, review
Spend against budget cap with forecast line and projected breach date, with the export control for a group-scoped view
The ceiling, the forecast, and the file finance asked for
A real use case, worked

The quarterly review of an AI support department.

A scenario: a support org runs a triage agent, two responder agents and an escalation employee against a 4-hour SLA. The quarter ends, finance and the ops lead sit down, and every number below is already in the ledger — nothing is reconstructed, surveyed or estimated.

Cost / outcome
What a resolved ticket costs
Because every message is token- and dollar-metered, the headline is a division, not a model: the quarter's spend over tickets resolved. And because the scorecard counts cost per successful run, retried and failed attempts are priced into the number instead of hidden under it.
Hours returned
What the humans got back
The run ledger shows what the department actually handled — every resolved ticket a person didn't touch, and every escalation a person did. Multiply the untouched volume by the team's own handle time and the hours-returned figure is the org's arithmetic on real counts, not a vendor's slide.
Reliability
Success rate and p95, per agent
The agent leaderboards rank the department's workers by success rate and duration. One responder sits mid-pack on volume but carries the worst failure rate; its scorecard names the cause from the failure-reason breakdown and shows the p95 creep against the SLA. That's next quarter's first fix, identified in the review itself.
Budget
Spend vs cap, with the forecast
The department's spend reads against the hard monthly cap set in Governance, with the period-end forecast and projected breach date in view. The conversation is 'raise the cap or tune the fleet' — held before the breach, with the breaker as the backstop, never after the invoice.
Chargeback
The export finance keeps
The quarter exports per group: each team's share of spend, attributed by the workers that drove it, with the date window and scope carried into the file. Finance books the chargeback from the export; the review ends with numbers owned, not disputed.
A group-scoped spend view for a quarter: per-group cost attribution with date window set, and the export control visible
The chargeback view: the quarter, attributed by group, ready to export
What you can answer

Six questions a vendor dashboard dodges.

What does one outcome cost?

Per-message metering rolls up to cost per successful run, per session, per completed task — division, not estimation.

Where is spend going?

The cost trend stacked by module plus Top-by-spend name the worker driving this month's creep.

Which agent fails most?

The highest-failure-rate leaderboard puts your worst worker at the top, ready to open and triage.

Which model should we run?

The Models table compares cost, latency and success per model from your own runs — evidence before you re-route traffic.

Is this employee earning its keep?

A falling intervention rate and a falling cost per completed task say the AI Employee is learning.

What does each team owe?

The Groups chargeback scorecard attributes org spend back to the team that drove it, and exports.

The payoff
100%
of runs costed natively: tokens and dollars per message, rolled up to cost per successful run.
~30 sec
to a full fleet-health read: spend drift, failures, budget risk, adoption, without a click.
6
entity types with their own leaderboards: agents, teams, employees, models, users, groups.
3
layers deep: fleet overview, entity leaderboard, single-worker scorecard — and every layer exports.
app.turtleaicoworker.com/statistics/models
The Models statistics screen: a performance table weighing each LLM by cost, latency and success rate, measured from the workspace's own runs
The Models screen: every LLM by cost, latency and success — measured from your traffic, not a vendor benchmark.
From the live catalog

Every pack reports here.

All templates
CA Firm: Compliance & Practice Pack
Solution Pack · Professional Services

Complete operating system for a Chartered Accountant practice. Tracks clients, engagements, statutory deadlines, and IT/GST notices. Includes AI agents for notice triage, GST reconciliation, filing reminders, and client communication. Comes with an AI Employee (Priya) who coordinates compliance work end-to-end.

4 agents5 tables1 employee
Official · ~15 min
AP and Bank Rec Desk
Solution Pack · Finance

The accounts payable and bank reconciliation desk a fractional controller runs for a client. Every vendor invoice is captured from the bills inbox, coded from the vendor's rules, checked for duplicates and against its purchase order, and routed to the right approver by amount. Approved invoices become a posting pack for the ledger and a weekly payment run. Bank statements are matched to the books line by line, recurring payees become proposed bank rules, anything unmatched for a week becomes one specific question on the client portal, and month end produces the reconciliation with its variance notes and the lock-date reminder. Comes with Cass, an AP and Close Coordinator who runs the queue.

8 agents11 tables1 employee
Official · ~30 min
Outbound Desk
Solution Pack · Sales

Signal-based outbound for a founder or a small GTM team. Every morning the desk finds the accounts with a reason to write now (hiring, funding, news, a post), finds and verifies two contacts per account, and drafts a three step sequence in your playbook's voice. You approve, and the email goes from your own inbox. Replies land in one queue, classified with the next step ready. Meetings get a one-page brief. Your CRM stays the record. Comes with Remy, an Outbound Coordinator who runs the morning queue.

9 agents9 tables1 employee
Official · ~25 min
Contracts and Signatures Desk
Solution Pack · Legal Ops

Contracts drafted, sent, chased, filed and watched, with a person at every step that matters. A colleague requests a contract from the portal; the drafter fills the approved template from the request and the CRM and marks what it could not fill. A person reviews and sends it for signature. Every morning the unsigned envelopes are chased and the old ones escalated. When a contract is signed its key terms (payment, liability, termination, renewal, governing law) are read from the PDF into a table with anything non-standard flagged, and the signed copy is filed by counterparty and type under your naming convention. Every Monday the contracts inside their notice window get a renew, renegotiate or terminate note for the owner. Works with Dropbox Sign, Google Docs, Google Drive and HubSpot. No agent ever signs, voids or counter-signs. Comes with Ren, a Contracts Coordinator, a Contracts Helpdesk and a portal for requesters.

5 agents7 tables1 employee
Official · ~30 min
Customer Support Desk
Solution Pack · Customer Support

The support team's queue, prepared. Every new ticket is read, categorised, given a priority and a drafted reply from your help centre within a minute; a person reads and sends. Questions that keep coming back become draft articles. Bugs are escalated to engineering with the steps and the evidence attached. Calls are summarised into a ticket. At the end of the day the team gets the volume, the response time and the three complaints of the day. Works with your helpdesk (Zendesk, Freshdesk, Intercom, Help Scout, Gorgias, Front and others), your phone tool and your issue tracker. Nothing is sent, refunded or closed by the software. Comes with Sol, a Support Coordinator.

6 agents6 tables1 employee
Official · ~25 min
Content Engine
Solution Pack · Marketing

Founder-led content for a founder or a small team. Every morning the desk listens (Reddit, LinkedIn, X, YouTube) for what your buyers are asking, turns the best of it into ideas by pillar, and drafts posts in your voice for LinkedIn and X. You approve; the desk hands you the final text to paste and post. Comments and reactions on what you published are harvested, and the people who match your ICP become leads. A newsletter issue is assembled from the week. Monday tells you which pillar and format earned attention. Comes with Theo, a Content Producer.

8 agents6 tables1 employee
Official · ~20 min
Questions

The details, up front.

Where do these numbers come from — and can we trust them in front of finance?
From the runs themselves. Every model call records tokens and dollars per message as it happens; every run totals cost, tokens and duration; every tool call is logged in the audit trail. The analytics are a read of that ledger — not a separate telemetry pipeline, not vendor-reported usage, not sampling. A quarterly figure decomposes all the way down to individual runs and messages if anyone asks.
Can we do chargeback by department?
Yes. The Groups screen attributes spend to the group that drove it, and any scoped view exports with its date window, granularity and scope intact — a group's quarter for chargeback, an agent's reliability for a post-mortem, a user's activity for a review. Chargeback becomes a filtered export, not a spreadsheet project.
What does "cost per successful run" mean, and why does it matter?
Most tools cost you per attempt. A worker's scorecard also counts cost per successful run: total spend divided by runs that actually succeeded. That's the number that catches a prompt change which quietly doubled your retries — per-attempt cost barely moves, cost per success jumps. It's also the number closest to what a CFO means by 'cost per outcome'.
Does everyone see the same numbers?
No. Every view is role-scoped through the same authorization matrix that governs the rest of the platform. Owners and Admins see org-wide health and cost, a Builder sees the assets they own, a Member sees a private view of their own activity. Fleet-wide analytics and org cost are Owner/Admin territory by design, not accident.
Can I compare models before switching?
Yes. The Models screen weighs each LLM's cost, latency and success rate side by side, measured from your own traffic rather than a vendor benchmark. Because models are chosen per agent, acting on the comparison is a settings change, not a rebuild — and the next period's numbers tell you whether it held.
How does this connect to budgets and governance?
Directly. Spend reads against the hard daily and monthly caps set in Governance, with a period-end forecast and projected breach date. Circuit breakers and budget warnings surface in the attention feed. The scorecard and the ceiling are two views of the same metered ledger — see the Governance page for the enforcement side.
The other half

The scorecard reads the ledger. Governance enforces it.

Governance