Platform · Measure · The ledger

Analytics

Six months into most AI rollouts, nobody can tell the CFO what it costs per outcome or what it replaced — finance wants chargeback, the board wants proof, and the vendor's dashboard shows API calls. Here every run is metered natively — tokens and dollars per message, per call, per worker — and rolls up into scorecards you can defend a budget with: cost per successful run, spend against a hard cap, exportable by group.

See your fleet's numbers Budgets & governance
At a glance
  • Meteringtokens + USD, per message
  • Cost accountingper successful run
  • Depthfleet → entity → run
  • Entity screens6 types
  • Budgetsforecast vs hard cap
  • Any viewrole-scoped · exportable
IN A SOLUTION PACK

In a pack, the ledger arrives wired: runs, cost, and outcomes tracked from the first install.

See Solution Packs →
app.turtlecoworker.com/statistics
The Statistics org overview: five KPI tiles with prior-period deltas, cost trend stacked by module, Top-by-spend leaderboard, activity heatmap and the attention feed
The org overview: five health KPIs, cost trend, spend leaderboard, activity heatmap, attention feed. One scroll, no clicks.
Why it's built this way

If you can't price the outcome,
you can't defend the budget.

AI spend without outcome accounting dies in the next budget cycle — not because it didn't work, but because nobody could prove it did. The pilot's champion is left arguing anecdotes against an invoice.

So the accounting is native, not bolted on. Every model call records its tokens and dollars per message on the way through; every run totals cost and duration; every tool call lands in the audit trail. Analytics is a read of that ledger — attributable down to the worker, the run and the message, and exportable in the shape finance asks for.

What you can measure

Eight lenses. One ledger.

From the metering underneath to the export finance takes away: every layer between a token and a defensible number.

01 · Metering

Costed at the message, so it's honest at the quarter.

Roll-up numbers are only as good as what's underneath, and most AI reporting is built on estimates. Here the accounting is native: every model call records its tokens and its dollars per message, every run totals tokens, cost and duration, and every tool call is classified and logged in the audit trail. The quarterly number is a sum of real rows, not a model of a model — which is why finance can lean on it.

  • Per-message accounting: tokens and USD recorded as the run happens
  • Per-run totals: tokens, cost, duration — joined to the audit trail
  • Nothing self-reported, nothing sampled: a sum, not an estimate
A single run's cost breakdown: per-message token counts and USD amounts rolling up to the run total
The atom of the ledger: one run, costed message by message
02 · Fleet health

The rollout's health, without a meeting.

Six months in, 'how is the AI going?' is usually answered with anecdotes. The org overview answers it in one scroll: five health KPIs with prior-period deltas, a spend trend, an activity heatmap, and a live attention feed of failures, tripped breakers, budget warnings and flagged PII. Whether spend is drifting, runs are failing or adoption is stalling is visible before anyone opens a ticket.

  • Five KPIs with Δ vs the prior period: success, spend, runs, adoption
  • A spend trend beside a top-by-spend leaderboard, so drift has a source
  • An attention feed: failures, tripped breakers, budget warnings, flagged PII
The org overview dashboard: KPI tiles with deltas, the spend trend, activity heatmap and attention feed in one scroll
One scroll, no clicks: the state of the rollout
03 · Spend

Cost drift gets a name before the invoice does.

The month the bill jumps, the vendor's answer is 'usage went up.' Here the spend trend sits beside a Top-by-spend leaderboard that ranks your most expensive workers. When the trend ticks up week-over-week, the worker driving it is one glance away — a specific agent, team or employee you can open, inspect and fix.

  • Daily spend with the workers behind it: agents, teams, employees
  • Top by spend ranks the most expensive workers, most costly first
  • From 'the bill went up' to the responsible worker in one click
shot: analytics-top-spendThe spend view: the spend trend beside the Top-by-spend leaderboard of named workers with their run counts and costs
The leaderboard that answers 'who spent it'
04 · Unit economics

Cost per outcome, not cost per attempt.

Cost per API call flatters everyone. A single worker's scorecard counts cost per successful run — total spend divided by the runs that actually finished the job — alongside a failure-reason breakdown, a p50/p95 duration trend and a trigger-source split. That distinction is where a prompt change that quietly doubled retries finally shows up: cost per attempt barely moves, cost per success jumps.

  • Cost per run and cost per successful run, side by side
  • Failure reasons named — a tool quota, a timeout — not guessed
  • p50/p95 duration trend and a trigger-source split
A single agent's scorecard: cost per run vs cost per successful run, failure-reason breakdown, p50/p95 duration trend
The honest number: what a finished job actually costs
05 · Compare

A leaderboard for every kind of worker.

Averages hide the worker that's quietly burning the budget. Entity Analytics ships one Statistics screen per type — Agents, Teams, AI Employees, Models, Users, Groups — each ranking its instances against each other. The Agents screen alone carries four leaderboards: most run, most expensive, highest failure rate, slowest. The interesting worker is the one whose rank differs across them: mid-pack on runs, first on failures.

  • Four agent leaderboards: most run · most expensive · highest failure · slowest
  • Six entity types, each ranked against its own kind
  • Bars turn warn/bad when a value crosses a threshold
The Agents statistics screen: four leaderboards ranking the same agents by runs, spend, failure rate and duration
The same agents, ranked four ways
06 · Teams & employees

Proof a team collaborates and an employee improves.

The Teams screen shows how each team actually performs — runs, cost and reliability per team, with member-level attribution. The AI Employees screen tracks whether a named employee is earning autonomy: a task funnel, a human-intervention rate that should be falling, and a memory panel that should be growing. A performance review becomes a chart, not an argument.

  • Per-team scorecards with member-level attribution
  • A task funnel: assigned, attempted, completed, escalated
  • A declining intervention rate as an employee earns autonomy
The Teams statistics screen: per-team scorecards with runs, cost and reliability for teams with real history
How the work moves, and whether the worker is learning
07 · Models

Migrate models on evidence, not a hunch.

Model choice is a recurring cost decision most teams make once and never revisit. The Models screen weighs every LLM you run by cost, latency and success rate, side by side, measured from your own traffic — not from a vendor benchmark. When a cheaper model would hold the line on an extraction step, the numbers say so, and the swap is a per-agent setting, not a rebuild.

  • A performance table: cost · latency · success, per model, from your runs
  • Evidence for a migration before you commit traffic to it
  • Catch the expensive model doing cheap work, and vice versa
The Models statistics screen: per-model table comparing cost, latency and success rate measured from the workspace's own runs
The migration case, made from your own traffic
08 · Budgets & export

Spend against a ceiling, and receipts that leave the app.

Numbers that stay in a dashboard don't survive a budget meeting. Spend here reads against the hard caps set in Governance — per workspace, employee or team, with a period-end forecast and the projected breach date in view. And any scoped view exports: a group's spend for chargeback, an agent's reliability for a post-mortem, a personal scorecard for a review. The date window, granularity and scope carry into the export.

  • Spend vs hard caps, with period-end forecast and breach date
  • Views are role-scoped: org-wide for Owners/Admins, own assets for Builders
  • Any scoped view exports — chargeback, post-mortem, review
Spend against budget cap with forecast line and projected breach date, with the export control for a group-scoped view
The ceiling, the forecast, and the file finance asked for
A real use case, worked

The quarterly review of an AI support department.

A scenario: a support org runs a triage agent, two responder agents and an escalation employee against a 4-hour SLA. The quarter ends, finance and the ops lead sit down, and every number below is already in the ledger — nothing is reconstructed, surveyed or estimated.

Cost / outcome
What a resolved ticket costs
Because every message is token- and dollar-metered, the headline is a division, not a model: the quarter's spend over tickets resolved. And because the scorecard counts cost per successful run, retried and failed attempts are priced into the number instead of hidden under it.
Hours returned
What the humans got back
The run ledger shows what the department actually handled — every resolved ticket a person didn't touch, and every escalation a person did. Multiply the untouched volume by the team's own handle time and the hours-returned figure is the org's arithmetic on real counts, not a vendor's slide.
Reliability
Success rate and p95, per agent
The agent leaderboards rank the department's workers by success rate and duration. One responder sits mid-pack on volume but carries the worst failure rate; its scorecard names the cause from the failure-reason breakdown and shows the p95 creep against the SLA. That's next quarter's first fix, identified in the review itself.
Budget
Spend vs cap, with the forecast
The department's spend reads against the hard monthly cap set in Governance, with the period-end forecast and projected breach date in view. The conversation is 'raise the cap or tune the fleet' — held before the breach, with the breaker as the backstop, never after the invoice.
Chargeback
The export finance keeps
The quarter exports per group: each team's share of spend, attributed by the workers that drove it, with the date window and scope carried into the file. Finance books the chargeback from the export; the review ends with numbers owned, not disputed.
A group-scoped spend view for a quarter: per-group cost attribution with date window set, and the export control visible
The chargeback view: the quarter, attributed by group, ready to export
What you can answer

Six questions a vendor dashboard dodges.

What does one outcome cost?

Per-message metering rolls up to cost per successful run, per session, per completed task — division, not estimation.

Where is spend going?

The cost trend stacked by module plus Top-by-spend name the worker driving this month's creep.

Which agent fails most?

The highest-failure-rate leaderboard puts your worst worker at the top, ready to open and triage.

Which model should we run?

The Models table compares cost, latency and success per model from your own runs — evidence before you re-route traffic.

Is this employee earning its keep?

A falling intervention rate and a falling cost per completed task say the AI Employee is learning.

What does each team owe?

The Groups chargeback scorecard attributes org spend back to the team that drove it, and exports.

The payoff
100%
of runs costed natively: tokens and dollars per message, rolled up to cost per successful run.
~30 sec
to a full fleet-health read: spend drift, failures, budget risk, adoption, without a click.
6
entity types with their own leaderboards: agents, teams, employees, models, users, groups.
3
layers deep: fleet overview, entity leaderboard, single-worker scorecard — and every layer exports.
app.turtlecoworker.com/statistics/models
The Models statistics screen: a performance table weighing each LLM by cost, latency and success rate, measured from the workspace's own runs
The Models screen: every LLM by cost, latency and success — measured from your traffic, not a vendor benchmark.
From the live catalog

Every pack reports here.

All templates
Procurement Ops Control Tower Pack
Solution Pack · Operations

End-to-end procurement operations solution pack with vendor registry, purchase request review, contract path assessment, embedded review team, and AI employee coordination.

3 agents3 tables1 employee
Official · ~35 min
CA Firm: Compliance & Practice Pack
Solution Pack · Professional Services

Complete operating system for a Chartered Accountant practice. Tracks clients, engagements, statutory deadlines, and IT/GST notices. Includes AI agents for notice triage, GST reconciliation, filing reminders, and client communication. Comes with an AI Employee (Priya) who coordinates compliance work end-to-end.

4 agents5 tables1 employee
Official · ~15 min
Recruiting and Staffing Ops
Solution Pack · Recruiting

Move candidates faster without cutting corners on fair hiring. Riley, your AI recruiting coordinator, screens each new candidate against the job's must-haves with must-have-by-must-have reasoning, coordinates interview scheduling by drafting availability requests, drafts honest candidate outreach and status updates, builds submittal packages for hiring managers or clients mapping experience to requirements, and gives the team a daily pipeline read. Screening never considers anything beyond skills and qualifications, and it never rejects a candidate or extends an offer; those stay human decisions. Everything else is draft-and-approve: candidate messages are prepared for a human to send, and interview times are proposed, never confirmed, by the agent. Four tables hold your job orders, candidates, interview scheduling, and pipeline metrics. A knowledge base holds your screening standards, submittal format, communication voice, and compliance guardrails, which every assessment and message follows. Automations included: new candidates are screened on arrival, plus an optional daily pipeline digest. Works out of the box with the roles and candidates you add; connect an ATS or job board later to sync candidates, and Gmail and Calendar to send outreach and confirm interviews from the platform. Best first step: replace the placeholder screening standards and compliance rules with your own and add an open job.

5 agents4 tables1 employee
Official · ~20 min
Sales Engine
Solution Pack · Sales

A complete outbound-to-pipeline sales engine for a small team. Define your ideal customer profile once in the ICP Profiles table and paste your product details into the Product Knowledge Base, then Sasha, your AI SDR, discovers and vets target accounts, enriches prospects, researches companies, scores leads, and drafts outreach, all grounded in your data. Inbound leads and your deal pipeline live in tables the founder can watch. Automations included: new prospects are auto-enriched on arrival, and an optional daily discovery run finds fresh accounts from your Active ICP. Requires an LLM provider plus the Apollo, Serper, Exa, and Google Search connectors. Replace the sample knowledge base content and the example ICP row with your own.

6 agents5 tables1 employee
Official · ~20 min
Customer Support Ops
Solution Pack · Support

Answer more support tickets, faster, without losing the human touch. Sam, your AI support specialist, classifies every incoming ticket by category, sentiment, and priority, drafts an accurate reply grounded only in your knowledge base, and escalates anything it is not confident about instead of guessing. It matches recurring problems to approved canned responses, mines new tickets into known issues so answers stay consistent, tracks the feature requests buried in tickets, and produces a weekly insights digest of volume, deflection, and top drivers. Everything is draft-and-approve: replies are prepared for a human to review and send, and sensitive tickets (anger, churn, refunds, account deletion, legal or privacy) are always escalated. Four tables hold tickets, known issues, feature requests, and daily metrics. A knowledge base holds your help-center content and support voice, which every reply cites. Automations included: a reply is drafted the moment a ticket arrives, plus an optional daily sweep of anything still unanswered. Works out of the box with no external tool; connect Gmail later to send from the platform and Slack for the digest. Best first step: replace the placeholder knowledge base with your real help-center articles or URLs.

5 agents4 tables1 employee
Official · ~20 min
Sales Engine Pro
Solution Pack · Sales

A complete, full-lifecycle sales engine for a growing team, spanning outbound prospecting and account management across one shared data layer, with a companion inbound concierge. Outbound: Sasha, your AI SDR, discovers and vets accounts, enriches prospects, scouts buying signals, researches companies, scores leads, and drafts cold email, LinkedIn, and nurture outreach. Account management: Maya, your AI Account Manager, qualifies opportunities, drafts and reviews proposals, flags pipeline risk, forecasts renewals and expansion, and keeps the CRM clean. Seven tables hold ICP profiles, target accounts, prospects, website leads, CRM contacts, opportunities, and customer accounts. Automations included: new prospects are auto-enriched on arrival, and an optional daily discovery run finds fresh accounts from your Active ICP. For the inbound engine, install the companion Website Sales Concierge team template, which routes visitor questions to product, pricing, and objection specialists and captures leads into the Website Leads table. Requires an LLM provider plus the Apollo, Serper, Exa, and Google Search connectors. Replace the sample knowledge base content and the example ICP row with your own. Outreach and proposals are drafted for a human to review and send.

15 agents7 tables2 employees
Official · ~25 min
Questions

The details, up front.

Where do these numbers come from — and can we trust them in front of finance?
From the runs themselves. Every model call records tokens and dollars per message as it happens; every run totals cost, tokens and duration; every tool call is logged in the audit trail. The analytics are a read of that ledger — not a separate telemetry pipeline, not vendor-reported usage, not sampling. A quarterly figure decomposes all the way down to individual runs and messages if anyone asks.
Can we do chargeback by department?
Yes. The Groups screen attributes spend to the group that drove it, and any scoped view exports with its date window, granularity and scope intact — a group's quarter for chargeback, an agent's reliability for a post-mortem, a user's activity for a review. Chargeback becomes a filtered export, not a spreadsheet project.
What does "cost per successful run" mean, and why does it matter?
Most tools cost you per attempt. A worker's scorecard also counts cost per successful run: total spend divided by runs that actually succeeded. That's the number that catches a prompt change which quietly doubled your retries — per-attempt cost barely moves, cost per success jumps. It's also the number closest to what a CFO means by 'cost per outcome'.
Does everyone see the same numbers?
No. Every view is role-scoped through the same authorization matrix that governs the rest of the platform. Owners and Admins see org-wide health and cost, a Builder sees the assets they own, a Member sees a private view of their own activity. Fleet-wide analytics and org cost are Owner/Admin territory by design, not accident.
Can I compare models before switching?
Yes. The Models screen weighs each LLM's cost, latency and success rate side by side, measured from your own traffic rather than a vendor benchmark. Because models are chosen per agent, acting on the comparison is a settings change, not a rebuild — and the next period's numbers tell you whether it held.
How does this connect to budgets and governance?
Directly. Spend reads against the hard daily and monthly caps set in Governance, with a period-end forecast and projected breach date. Circuit breakers and budget warnings surface in the attention feed. The scorecard and the ceiling are two views of the same metered ledger — see the Governance page for the enforcement side.
The other half

The scorecard reads the ledger. Governance enforces it.

Governance