Discover

Find patterns you didn’t know to look for

Go from raw production data to actionable insights with AI-powered analysis. Topics surfaces patterns automatically; Loop lets teams explore data through natural language.

Observability tells you what happened. Discovery tells you what it means.

When your agents handle thousands of requests daily, you can't read every trace. You need tools to surface failure modes and edge cases without requiring you to know what to look for.

With discovery, patterns surface and you can act on evidence. Topics distills thousands of logs into themed clusters you can use immediately.

Two tools for making sense at scale

From individual logs to actionable insight

Topics automatically clusters your logs into themes like user intents, failure modes, and edge cases, without you defining categories upfront. Loop lets you ask follow-up questions in plain language and get answers backed by your production data.

Topics

AI-powered clustering

Groups traces by semantic similarity. UMAP + HDBSCAN + c-TF-IDF.

Task
User intent or goal
24.0%203 items
Refund requests

Customers are asking for money back after missing deliveries, delayed refund status updates, and policy confusion around damaged or returned orders.

18.4%156 items
Sizing and fit questions

Shoppers need help choosing sizes, comparing fit across products, and deciding whether an item will work before placing or exchanging an order.

15.0%127 items
Order tracking

Users want current shipment status, tracking links, delivery estimates, and explanations for packages that appear stalled or delayed in transit.

10.9%92 items
Damaged items

Customers report products arriving broken, packaging failures, and requests for replacements where the agent needs to gather evidence and route next steps.

9.0%76 items
Subscription questions

Users ask how renewals, cancellations, and subscription refunds work, often needing policy-grounded answers before support escalation.

Loop

Ask your data anything

Natural language queries against your logs.

Loop agent
What are the common failure modes of my support agent?
Generated search plan
Processed traces847
SearchedRefund policy failures

Top failure mode: subscription refunds. 18 of 20 low-scoring traces involve refund requests where the agent skips policy lookup.

Want me to bootstrap a scorer or generate a regression dataset from these traces?

Created LLM judge scorer
Ask questions about logs
Powered by Brainstore

Built for agent data at scale

Discovery only works if queries stay fast at production scale. Brainstore is Braintrust's database for agent observability. Search and filter millions of traces in under a second, including full-text search across prompts and error messages.

0.0x
Faster full-text search
Competition
0 ms
Brainstore
0 ms
0.00x
Faster write latency
Competition
0 ms
Brainstore
0 ms
0.00x
Faster span load time
Competition
0 ms
Brainstore
0 ms
Learn more about Brainstore
The automation suite

Evals and observability, automated

Discovery surfaces what to fix. Automation makes sure it stays fixed.

Latency

Alerts

Notifications when scores, latency, or errors cross boundaries.

Set up alerts
Created
Input
Factuality
Helpfulness
Sentiment
Jan 25, 2:34 PM
Hi! What's the status of order #12345?
94%
91%
82%
Jan 25, 2:31 PM
Can I return my order from last week?
91%
88%
75%
Jan 25, 2:28 PM
What are your shipping rates to Canada?
88%
86%
79%
Jan 25, 2:25 PM
I need to change my delivery address
96%
93%
85%
Jan 25, 2:22 PM
Do you have any discounts for bulk orders?
89%
87%
88%
Jan 25, 2:19 PM
My package arrived damaged
92%
84%
45%
Jan 25, 2:16 PM
What payment methods do you accept?
97%
92%
81%
Jan 25, 2:08 PM
Where is my refund? It's been 2 weeks!
78%
55%
32%
Jan 25, 2:01 PM
How do I track my order?
95%
90%
77%
Jan 25, 1:58 PM
Can I pick up my order in store?
93%
89%
84%
Jan 25, 1:55 PM
When will order #88219 ship?
91%
87%
71%

Online scoring

Production traffic scored continuously. Same scorer library.

Score production traffic
braintrust-evalBotcommented1 hour ago• edited
Support agent (HEAD-1714341466)
ScoreAverageImprovementsRegressions
Policy compliance94% (+4%)12 🟢2 🔴
Resolution quality87% (+6%)18 🟢5 🔴
Escalation accuracy91% (+2%)9 🟢3 🔴
Duration1.8s (-0.4s)11 🟢2 🔴
Review in Braintrust

Quality gates

Block deployments when eval scores drop below thresholds.

Add to CI
Isolated cloud sandbox
AWSModal

Sandboxed evals

Push eval code once. Teammates run complex agents from the playground without local setup.

Run sandbox evals
Customer stories

From reactive to proactive

Sarah Sachs, AI Lead

“There are some problems we wouldn't know were problems without Braintrust.”

Luis Héctor Chávez, CTO

“Braintrust helped us identify several patterns that we wouldn't have found.”

Allen Kleiner, AI Engineering Lead

“Loop helps us understand trace details that would be impossible to scan manually.”

Start building

Create a free account and start monitoring in minutes.