<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Distributional Blog</title><description>Company news and the Distributional technical archive, preserved at its original addresses.</description><link>https://distributional.com</link><language>en-us</language><item><title>Distributional&apos;s next chapter: Talaria Scientific</title><link>https://distributional.com/blog/distributional-is-now-talaria</link><guid isPermaLink="true">https://distributional.com/blog/distributional-is-now-talaria</guid><description>Distributional is now Talaria Scientific: same company, same investors, hard pivot. The honest story of what worked, what didn&apos;t, and the new mission one level deeper.</description><pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate><content:encoded>
&lt;p&gt;Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. Same company, same investors, hard pivot, new mission. This post is the honest version of how we got here and where we&amp;#39;re going.&lt;/p&gt;
&lt;h2 id=&quot;what-we-set-out-to-do&quot;&gt;&lt;a href=&quot;#what-we-set-out-to-do&quot;&gt;What we set out to do&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;We started Distributional in 2023 to solve AI testing for the enterprise: give teams a rigorous, repeatable way to know whether the AI systems they were shipping actually behaved the way they believed by using Bayesian statistical tests to detect non-stationarity and chaotic behavior. In late 2025 we sharpened that into behavioral analytics for AI agents in production, finding the patterns in raw trace data to find the evals that slipped through the cracks.&lt;/p&gt;
&lt;p&gt;The vision landed, more than once. Enterprises took meetings and ran pilots, partners brought us opportunities, and inbound arrived steadily. The interest was real, but we never got the organic, in-production pull that separates a product from a promising idea. We had vision-market-fit more than once, but never turned it into product-market-fit while the market fundamentally changed underneath us.&lt;/p&gt;
&lt;h2 id=&quot;what-worked-and-what-didnt&quot;&gt;&lt;a href=&quot;#what-worked-and-what-didnt&quot;&gt;What worked and what didn&amp;#39;t&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;I still believe better AI testing is obvious and inevitable, but we failed to deliver the product the market needed at the time. The market always wins and being too early is the same as being wrong. Fundamentally, I think it broke down into three main issues. First, people don&amp;#39;t like tests that slow them down, even (or especially!) if they don&amp;#39;t write those tests themselves. People especially don&amp;#39;t like complex, high dimensional statistical tests that are hard to interpret and take action on (you can&amp;#39;t change the weights!). Finally, the market decided it would just YOLO AI systems into production anyway; a statistical gate between them and shipping was a hard sell. That last point was especially surprising to me after watching AI teams struggle to adopt basic ML models in my SigOpt days because random forests and gradient boosted decision trees were too exotic (how times have changed!).&lt;/p&gt;
&lt;p&gt;I also believe that AI analytics is going to be an extremely important part of the stack. Analytics helps you find the evals that you don&amp;#39;t know to look for by finding patterns in your unstructured logs. I think it is a core component of &lt;a href=&quot;https://twimlai.com/podcast/twimlai/how-find-agent-failures-your-evals-miss&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Maslow&amp;#39;s hierarchy of observability&lt;/a&gt;, like any major category before AI, analytics helps you find the unknown unknowns that help make your product better over time. But it didn&amp;#39;t work as a product for a different reason: it is a feature, not a product, let alone a startup. We were also selling from the top of the hierarchy while most of our buyers were still building the foundation underneath it. It was too easy for a monitoring product to box us out with simple analytics on top of their solutions, and most customers were just trying to stand up any agent, not improve them over time. Behavioral analytics is what you want once logging, tracing, and evals are in place, and much of the market wasn&amp;#39;t there yet.&lt;/p&gt;
&lt;h2 id=&quot;the-decision&quot;&gt;&lt;a href=&quot;#the-decision&quot;&gt;The decision&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This spring we ran a thorough process to find the company a home. After thousands of hours of work we came close to a good outcome: serious diligence, term sheets on the table, and ultimately an acquihire offer for the whole team.&lt;/p&gt;
&lt;p&gt;But none of those paths was better than the remaining one: a hard pivot. Keep the company, keep the balance sheet, and point everything we had learned at a problem where we have an unfair advantage.&lt;/p&gt;
&lt;p&gt;Taking care of the team was extremely important to me and I&amp;#39;m proud that nearly the whole team received offers as part of our M&amp;amp;A search, and most are landing together at their next adventure. None of this part was easy, and the people involved handled it with professionalism and grace. I hope I can work with them again in the future.&lt;/p&gt;
&lt;h2 id=&quot;same-company-new-mountain&quot;&gt;&lt;a href=&quot;#same-company-new-mountain&quot;&gt;Same company, new mountain&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Talaria Scientific is a new product direction for the same Delaware corporation, Distributional, backed by the same investors. This is a change of mission, not a spinout or a new company. We&amp;#39;re building a multi-agent harness for computational science: tooling that turns a 10x researcher into a 100x researcher, the way agentic coding did for software engineers over the last three years. You can read more about what we are building &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I truly believe that the path to abundance via AI isn&amp;#39;t in replacing white collar jobs, it is about new scientific discoveries that lead to things like new materials and more efficient energy which in turn lead to advances in medicine, aerospace, and more. It is about living in a scifi future and improving quality of life for everyone.&lt;/p&gt;
&lt;p&gt;We can now meaningfully accelerate science with AI. And seen from far enough back, the new mission is the fourth incarnation of a process I have been on my entire career: make expensive computational work optimal, scalable, and trustworthy. Talaria takes that work one level deeper, from the tooling around the science to accelerating the science itself.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://distributional.com/images/blog/distributional-is-now-talaria/figure.png&quot; alt=&quot;One mission, four eras: optimize expensive computational work, run it at scale, trust it in production, and now point it at science itself. The box around the last two is the same-entity part: Distributional and Talaria Scientific are one company with the same investors. The mission is what moved one level deeper.&quot; /&gt;&lt;figcaption&gt;One mission, four eras: optimize expensive computational work, run it at scale, trust it in production, and now point it at science itself. The box around the last two is the same-entity part: Distributional and Talaria Scientific are one company with the same investors. The mission is what moved one level deeper.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The slow, expensive part of research is the undifferentiated work in data wrangling, sysadmin, and validation. We aim to make that part fast, with the scientist in the loop and HPC as the guardrail. It&amp;#39;s a mech suit for the researcher, not an android auto-scientist that replaces them. The product will be built in the open at &lt;a href=&quot;https://talariasci.com/blog&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;talariasci.com&lt;/a&gt; and open source by &lt;a href=&quot;https://neurips.cc/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NeurIPS 2026&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Why do I believe this one is different? I&amp;#39;m the target audience, so I can dogfood it on my own research and learn in days what enterprise sales cycles taught me in quarters. It sits on the parts of my background that are hardest to copy: optimization research, HPC, and the scar tissue from trying to optimize it all (ie building evals for the last 15+ years). And the need is real: researchers feel this pain every day, and the software world just spent three years proving the model for the SWE workflow.&lt;/p&gt;
&lt;h2 id=&quot;what-happens-to-this-site&quot;&gt;&lt;a href=&quot;#what-happens-to-this-site&quot;&gt;What happens to this site&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;distributional.com stays up. The technical writing this team published deserves a stable home, so we&amp;#39;re restoring a select blog archive at its original addresses over the coming weeks, and corporate updates like this one will land here when there&amp;#39;s something real to say. Day to day, the building now happens at &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;talariasci.com&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;thank-you&quot;&gt;&lt;a href=&quot;#thank-you&quot;&gt;Thank you&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;To the team: you built something I remain proud of, and you handled the hardest parts of the last few years with grace. To our customers and partners: thank you for the trust and candor, especially the hard truths that led us to failing fast and getting to a better place. To our investors: thank you for continuing to believe and letting us continue to fight to make a difference in the world.&lt;/p&gt;
&lt;p&gt;We set out to make AI trustworthy enough for serious work. I still think that&amp;#39;s one of the defining problems of this era, and I suspect someone will crack the business of it. We&amp;#39;re going to keep working on it one level deeper: making the harness that makes computational science trustworthy, and fast.&lt;/p&gt;
&lt;p&gt;It&amp;#39;s time to keep building. Semper Deinceps.&lt;/p&gt;
&lt;p&gt;Scott&lt;/p&gt;</content:encoded><category>news</category></item><item><title>Beyond averages: how DBNL&apos;s distribution comparison reveals what summary metrics hide</title><link>https://distributional.com/blog/beyond-averages-how-dbnls-distribution-comparison-reveals-what-summary-metrics-hide</link><guid isPermaLink="true">https://distributional.com/blog/beyond-averages-how-dbnls-distribution-comparison-reveals-what-summary-metrics-hide</guid><description>Why P95s and averages hide bimodal regressions: four use cases for distribution comparison and the design decisions behind them. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Your P95 latency doubled after a deployment. Your average cost-per-call looks fine. A dashboard of aggregate numbers tells you &lt;em&gt;something changed&lt;/em&gt; — but not &lt;em&gt;what changed&lt;/em&gt;, or &lt;em&gt;for whom&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;This is the gap between metrics and understanding. Summary statistics compress the shape of your data into a single number, and in doing so, they erase the patterns that actually explain what&amp;#39;s happening. A bimodal latency spike — where half your requests are fast and the other half are stuck — looks identical to a uniform slowdown when reduced to a P95. A long tail of expensive LLM calls disappears into a healthy-looking average.&lt;/p&gt;
&lt;p&gt;DBNL&amp;#39;s Distribution tab was built to close that gap. It lives inside the Explorer — the primary analytical surface in DBNL&amp;#39;s observability platform, where engineers and data teams investigate trends, compare distributions, and trace agent execution paths across their LLM-powered applications. The Distribution tab is where you go when you need to move beyond trend lines and summary stats: it lets you visualize the full shape of any metric across your agent traces and compare that shape between two filtered populations, side by side on a shared axis. This post walks through why distribution analysis matters for LLM agent observability, the real-world use cases it unlocks, and the design decisions that make it work.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=xo_5nLR4D-I&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;why-distributions-matter-for-llm-agent-analytics&quot;&gt;&lt;a href=&quot;#why-distributions-matter-for-llm-agent-analytics&quot;&gt;Why distributions matter for LLM agent analytics&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;LLM-powered agents are inherently non-deterministic. The same prompt can produce wildly different execution paths, token counts, latencies, and costs depending on model behavior, tool-call branching, and context length. This makes aggregate statistics especially unreliable — a mean or percentile can mask the multimodal behavior that agents routinely produce.&lt;/p&gt;
&lt;p&gt;Distribution analysis answers the questions that summary metrics can&amp;#39;t:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Is this a uniform shift or a population split?&lt;/strong&gt; When latency increases, a distribution reveals whether all requests slowed down or whether a subset is stuck while the rest are fine. This distinction changes the debugging approach entirely — a uniform shift points to infrastructure, while a bimodal split points to specific code paths or input characteristics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Where does the mass actually sit?&lt;/strong&gt; Knowing that P95 latency is 8 seconds tells you nothing about the other 94%. A histogram shows whether most requests cluster tightly around 200ms with a thin tail, or whether the distribution is flat and unpredictable. The former is a healthy system with occasional outliers; the latter is a system with a reliability problem.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How do two populations actually differ?&lt;/strong&gt; Comparing &amp;quot;error traces vs. successful traces&amp;quot; or &amp;quot;experiment A vs. experiment B&amp;quot; with summary stats gives you two numbers. Comparing their distributions shows you &lt;em&gt;where&lt;/em&gt; they diverge — maybe errors only occur in a specific latency band, or maybe the cost distributions are identical except for a spike at the upper extreme.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;use-cases-when-to-reach-for-distribution-comparison&quot;&gt;&lt;a href=&quot;#use-cases-when-to-reach-for-distribution-comparison&quot;&gt;Use cases: when to reach for distribution comparison&lt;/a&gt;&lt;/h2&gt;
&lt;h3 id=&quot;debugging-production-regressions&quot;&gt;&lt;a href=&quot;#debugging-production-regressions&quot;&gt;Debugging production regressions&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Your alerting fires on a P95 latency spike. You open the Distribution tab, select &lt;code&gt;duration_ms&lt;/code&gt; at the trace level, and create two segments:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Segment A&lt;/strong&gt;: &lt;code&gt;status = error&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Segment B&lt;/strong&gt;: &lt;code&gt;status = success&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The histogram immediately shows the story. Successful traces cluster in a tight band around 300ms. Error traces show a completely separate mode at 12–15 seconds — they&amp;#39;re not just slow, they&amp;#39;re hitting a timeout boundary. The regression isn&amp;#39;t a uniform degradation; it&amp;#39;s a specific failure path that&amp;#39;s pulling the P95 up. You now know exactly which traces to drill into.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=HcK_IN-nHwg&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;evaluating-experiment-variants&quot;&gt;&lt;a href=&quot;#evaluating-experiment-variants&quot;&gt;Evaluating experiment variants&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;You&amp;#39;ve shipped a new prompt template optimized for cost efficiency — the Trends tab confirms average &lt;code&gt;`total_cost`&lt;/code&gt; dropped 8%. But cheaper doesn&amp;#39;t mean better. Did the leaner prompt sacrifice output quality? In the Distribution tab, you select &lt;code&gt;`output_relevancy`&lt;/code&gt; — an LLM-as-judge metric that classifies each response as &amp;quot;relevant&amp;quot; or &amp;quot;irrelevant&amp;quot; — and create two segments using experiment filters: - &lt;strong&gt;Segment A&lt;/strong&gt;: Experiment variant &lt;code&gt;`control`&lt;/code&gt; - &lt;strong&gt;Segment B&lt;/strong&gt;: Experiment variant &lt;code&gt;`new-prompt-v2`&lt;/code&gt; The histogram makes the trade-off immediately visible. In Relative mode, the control prompt produces irrelevant responses roughly 10% of the time. The new prompt doubles that to 20%. The cost savings are real, but they come at the expense of output quality for one in five requests. This is the kind of regression that an average relevancy score won&amp;#39;t catch — it&amp;#39;s only visible when you compare the full distribution of outcomes side by side. DBNL&amp;#39;s filter builder supports experiment-level filters natively, so setting up this comparison takes seconds. You select the experiment, pick the variants, and the histogram updates immediately — no need to export data, write SQL, or build a notebook. &lt;a href=&quot;https://www.youtube.com/watch?v=ZIwvSWGY-lM&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Watch on YouTube&lt;/a&gt; ### Comparing model providers or versions You&amp;#39;re evaluating whether to migrate from one model version to another. Beyond just checking average quality scores, you want to understand the full distribution of your custom evaluation metrics. Set up two segments filtered by model version, select your LLM-as-judge score column, and toggle to Relative mode. Even if one model version has 10x more traffic, the normalized view lets you compare distribution shapes directly. You might find that the new model has a tighter distribution (more consistent quality) even if the averages are similar — or that it eliminates the low-quality tail that was driving user complaints. &lt;img src=&quot;https://distributional.com/images/blog/beyond-averages-how-dbnls-distribution-comparison-reveals-what-summary-metrics-hide/69ea819f9550ab0fec0032f5_image1.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt; ### Validating guardrails and thresholds You&amp;#39;ve set a 45-second timeout on agent execution. How much traffic is actually approaching that boundary? View the &lt;code&gt;`duration_ms`&lt;/code&gt; distribution across all traces, then add a segment filtered to &lt;code&gt;`duration_ms &amp;gt; 40000`&lt;/code&gt;. The relative mass of the near-timeout population tells you whether the guardrail is an edge case or a significant portion of your traffic. If 12% of traces are within 5 seconds of the timeout, you have a systemic issue — not an outlier problem. ### Understanding user cohort behavior Do power users and new users drive different workload profiles? Filter your session-level &lt;code&gt;`duration_ms`&lt;/code&gt; distribution by a user cohort attribute and compare. You might discover that new users have short, simple sessions that cluster tightly, while power users produce a broad, flat distribution — indicating they&amp;#39;re pushing the agent into diverse, long-running workflows that need separate performance optimization.&lt;/p&gt;
&lt;h2 id=&quot;design-decisions-why-dbnls-distribution-tab-works&quot;&gt;&lt;a href=&quot;#design-decisions-why-dbnls-distribution-tab-works&quot;&gt;Design decisions: why DBNL&amp;#39;s distribution tab works&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Building a distribution comparison tool sounds straightforward — render a histogram, add a filter. In practice, most implementations fall apart when they try to generalize. Here&amp;#39;s how we approached the hard design problems.&lt;/p&gt;
&lt;h3 id=&quot;one-column-varied-filters--not-varied-columns&quot;&gt;&lt;a href=&quot;#one-column-varied-filters--not-varied-columns&quot;&gt;One column, varied filters — not varied columns&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;An early design considered letting users compare distributions of different columns side by side — latency vs. cost, for example. We ruled this out intentionally.&lt;/p&gt;
&lt;p&gt;Different columns can have fundamentally different data types and scales. A numeric column like &lt;code&gt;duration_ms&lt;/code&gt; produces a continuous histogram with range-based bins. A categorical column like &lt;code&gt;status&lt;/code&gt; produces a bar chart with discrete labels. Overlaying these on a shared axis is meaningless: the X-axis semantics don&amp;#39;t align, the bin boundaries are incompatible, and the Y-axis scales can differ by orders of magnitude. Even comparing two numeric columns breaks down when one ranges from 0–100ms and another from $0–$50.&lt;/p&gt;
&lt;p&gt;Instead, the comparison axis is &lt;em&gt;filters on the same column&lt;/em&gt;. This guarantees both segments share the same data type, the same bin boundaries, and the same axis scale. The question shifts from &amp;quot;how does column X compare to column Y?&amp;quot; to &amp;quot;how does column X behave under condition A vs. condition B?&amp;quot; — a question that only a shared-axis histogram can answer clearly.&lt;/p&gt;
&lt;p&gt;This constraint doubles as a guardrail: every valid input produces a coherent chart. Users can&amp;#39;t accidentally build a visualization that&amp;#39;s broken or misleading.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/beyond-averages-how-dbnls-distribution-comparison-reveals-what-summary-metrics-hide/69ea824bb7587ed902419e33_image2.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h3 id=&quot;two-segments-not-more&quot;&gt;&lt;a href=&quot;#two-segments-not-more&quot;&gt;Two segments, not more&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;We capped the segment limit at two. This was a deliberate constraint, not a technical limitation.&lt;/p&gt;
&lt;p&gt;Histograms are fundamentally different from line charts. A line chart can layer many series legibly because lines occupy different vertical positions and the eye tracks each one independently. Histogram bars compete for the same horizontal space. With three or more segments, bars within each bin become too narrow to read, hover targets shrink below useful thresholds, and the cognitive load of tracking color mappings across dozens of thin slices outweighs the analytical value.&lt;/p&gt;
&lt;p&gt;Two segments cover the dominant use case — A/B comparison. &amp;quot;How does this metric look for errors vs. non-errors?&amp;quot; &amp;quot;Production vs. staging?&amp;quot; &amp;quot;Experiment A vs. B?&amp;quot; — these are the questions that distribution comparison exists to answer, and all of them are two-population problems.&lt;/p&gt;
&lt;h4 id=&quot;grouped-bars-not-overlaid-or-separated&quot;&gt;&lt;a href=&quot;#grouped-bars-not-overlaid-or-separated&quot;&gt;Grouped bars, not overlaid or separated&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;With two segments to display, there are three reasonable bar layout approaches. We evaluated all three:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Overlaid (semi-transparent bars stacked on top of each other)&lt;/strong&gt; is common in statistical tools. It works when distributions differ significantly, but when they overlap heavily — the common case in A/B comparisons — the colors blend and individual values become impossible to read. Worse, hover and click targeting becomes ambiguous. In a drilldown-heavy interface where clicking a bar navigates to filtered logs, ambiguous click targets are a non-starter.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fully separated (two side-by-side charts)&lt;/strong&gt; eliminates overlap but makes direct comparison harder. The user&amp;#39;s eye has to travel between charts and mentally align bin positions. Subtle shape differences become invisible. Each chart is also compressed to half the available width, reducing visual resolution. It turns a comparison task into a memory task — hold the shape of one chart in mind while looking at the other — which is exactly what good visualization should eliminate.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Grouped (bars side-by-side within each bin)&lt;/strong&gt; is what we chose. Each bin gets a pink bar (Segment A) and an orange bar (Segment B) sitting next to each other. Spatial proximity makes per-bin comparison instant. Click targets are unambiguous. The distribution shapes remain visually intact because each segment&amp;#39;s bars form a continuous silhouette. The trade-off is that individual bars are narrower than in a single-segment view, but with bin count controlled by the backend, this stays readable.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&quot;count-vs-relative-two-lenses-on-the-same-data&quot;&gt;&lt;a href=&quot;#count-vs-relative-two-lenses-on-the-same-data&quot;&gt;Count vs. relative: two lenses on the same data&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;The chart toolbar provides a Count/Relative toggle that switches the Y-axis between raw counts and normalized ratios. This is essential for comparing segments of different sizes.&lt;/p&gt;
&lt;p&gt;If Segment A has 50,000 traces and Segment B has 5,000, count mode will show Segment A visually dominating every bin — making shape comparison impossible. Relative mode normalizes each segment independently (each sums to 1), so the chart shows &lt;em&gt;shape&lt;/em&gt; rather than &lt;em&gt;volume&lt;/em&gt;. Two bars of similar height in relative mode confirm the distributions have the same shape. Height discrepancies in specific bins reveal exactly where they diverge.&lt;/p&gt;
&lt;p&gt;The grouped bar layout makes this toggle especially effective. In count mode, you see volume differences. In relative mode, you see shape differences. Both comparisons are legible with side-by-side bars in a way that overlay cannot support.&lt;/p&gt;
&lt;h2 id=&quot;the-bigger-picture&quot;&gt;&lt;a href=&quot;#the-bigger-picture&quot;&gt;The bigger picture&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distribution comparison isn&amp;#39;t a standalone feature — it&amp;#39;s part of DBNL&amp;#39;s connected investigation workflow. You can start with a trend line that shows a metric changing over time, click a data point to drill into the distribution for that specific time window, and then click a histogram bar to see the raw traces that fall in that bucket. Each step narrows the investigation, and each transition preserves context.&lt;/p&gt;
&lt;p&gt;This is what purpose-built LLM agent analytics looks like. Not a generic charting tool with an AI label, but an opinionated system designed around the specific questions that agent developers ask: &lt;em&gt;Is my agent reliable? Is it cost-efficient? Did this change make it better or worse — and for which users?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The Distribution tab answers the version of those questions that summary metrics can&amp;#39;t reach. And in a world where agent behavior is inherently variable, reaching past the averages isn&amp;#39;t optional — it&amp;#39;s where the real understanding begins.&lt;/p&gt;
&lt;h2 id=&quot;get-started&quot;&gt;&lt;a href=&quot;#get-started&quot;&gt;Get started&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;You can get started with these new features in less than 20 minutes with our free, open, and installable sandbox at &lt;a href=&quot;https://docs.dbnl.com/get-started/quickstart&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://docs.dbnl.com/get-started/quickstart&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;. Once you are ready, you can also install our free, open full service in your own environment at &lt;a href=&quot;https://docs.dbnl.com/platform/deployment&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://docs.dbnl.com/platform/deployment&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;We are always happy to learn more about your use case and enterprise needs, so reach out to &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;contact@distributional.com&lt;/a&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Examples of issues that break agents in production</title><link>https://distributional.com/blog/examples-of-issues-that-break-agents-in-production</link><guid isPermaLink="true">https://distributional.com/blog/examples-of-issues-that-break-agents-in-production</guid><description>Concrete failure modes analytics catches in production agents: evals that pass but do not generalize, topic shifts, and complexity-correlated issues. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;As we discuss how our production agent analytics work, one of the first questions we normally get is, “Will you give me a concrete example?” This is because we are usually distinguishing between monitoring and analytics.&lt;/p&gt;
&lt;p&gt;Monitoring is useful for assessing thresholds on known issues or opportunities or evals in near real-time. Analytics are useful for finding signals on unknown issues or opportunities that are hidden in production agent traces. AI product teams use analytics to know what to monitor.&lt;/p&gt;
&lt;p&gt;In this post, we’ll share a few examples of issues that analytics would help identify. These issues were unknown prior to this analysis, are relatively nuanced correlations across multiple metrics, or some combination of both. The goal is to give you more intuition on what you’d get from using DBNL on your production traces.&lt;/p&gt;
&lt;p&gt;This post was heavily inspired by a conversation between Scott Clark, Distributional co-founder and CEO, and Jason Liu of the Developer Experience team at OpenAI. As &lt;a href=&quot;https://www.linkedin.com/in/sc932/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Scott&lt;/a&gt; put it in his discussion with &lt;a href=&quot;http://linkedin.com/in/jxnlco&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Jason Liu&lt;/a&gt;, “If you don’t run the test for a disease, it doesn’t mean you don’t have it.” Analytics is about finding those issues (or opportunities) you don’t know you have. You can watch this here (especially minutes 30-40):&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=QBaSnETJr4Q&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;examples-of-production-issues-that-analytics-can-catch&quot;&gt;&lt;a href=&quot;#examples-of-production-issues-that-analytics-can-catch&quot;&gt;Examples of production issues that analytics can catch&lt;/a&gt;&lt;/h2&gt;
&lt;h3 id=&quot;offline-pre-deployment-evals-that-fail-to-perform-out-of-sample-in-production&quot;&gt;&lt;a href=&quot;#offline-pre-deployment-evals-that-fail-to-perform-out-of-sample-in-production&quot;&gt;&lt;strong&gt;Offline, pre-deployment evals that fail to perform out of sample in production&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Your agent gets a 100% pass rate on an eval, but this eval is a kindergarten level math test. When users ask questions beyond kindergarten-level math, it underperforms.&lt;/li&gt;
&lt;li&gt;Your agent is evaluated on the rate at which it catches and passes negativity in an input to a specific guardrail. When it scales in production, however, it also tends to catch misspellings that are classified as negative, which sends far too many inputs to the guardrail, creating a poor user experience.&lt;/li&gt;
&lt;li&gt;You design your agent around a concentrated group of users in a few countries. Then you scale your agent internationally, and discover edge cases that result in degradations in performance that weren’t contemplated with pre-production evals.&lt;/li&gt;
&lt;li&gt;You correlate low feedback scores with topics and find that there are a subset of tasks the agent fails to perform well. You re-engineer the tools to account for this, and add these tasks to your eval set for future development and monitoring.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;insights-in-production-on-user-topics-intent-and-input-patterns&quot;&gt;&lt;a href=&quot;#insights-in-production-on-user-topics-intent-and-input-patterns&quot;&gt;&lt;strong&gt;Insights in production on user topics, intent, and input patterns&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;You’ve scaled an internal multi-turn chat RAG application across multiple functional areas (e.g., HR, IT, finance). Now you lack visibility into how people are using this system. Without insights on topics, it is hard to add additional support or to evaluate how the agent is performing on those queries.&lt;/li&gt;
&lt;li&gt;Your agent goes viral in Turkey and the majority of prompts are now in Turkish. Because you wrote the system prompt in English, however, it creates a bug where half the time the agent responds in Turkish, but half the time it responds in English, creating a bad user experience.&lt;/li&gt;
&lt;li&gt;You launch an agent with a marketing campaign targeting developers as the expected user base. Instead, managers find the agent more valuable and start to use it more than developers. Due to this shift, you need to rework prompt, context, and tools to cater to this different manager-level user base.&lt;/li&gt;
&lt;li&gt;You build an agent for productivity software workflows based largely on the assumption that email was the most important medium. This includes a lot of work iterating on email search tooling to make sure this is a strong aspect of the product. In production, you discover that more than 30% of search queries were based on photos – someone snapping a shot of a receipt, etc. – and not email. The agent is not designed to perform in this scenario, resulting in a poor user experience.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;issues-in-the-complexity-of-agent-behavior&quot;&gt;&lt;a href=&quot;#issues-in-the-complexity-of-agent-behavior&quot;&gt;&lt;strong&gt;Issues in the complexity of agent behavior&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;You develop an agent with a variety of LLM calls and access to a diverse set of tools. You organize the agent in a workflow with a relatively structured framework. When you push to production, you notice 403s when there is a specific prompt that results in the agent calling a specific tool. You notice you didn’t include a retry step in that process and it is consistently failing the first time, driving the error.&lt;/li&gt;
&lt;li&gt;You correlate tool calls, token cost, topics, and response quality metrics, and discover that there is a specific topic where the agent is inefficiently calling a tool that is too expensive.&lt;/li&gt;
&lt;li&gt;You find a spike in user frustration. Once you correlate this with topics, you discover that users are frustrated because their queries are being routed to a guardrail when they should actually be answered. You redesign the guardrail to avoid overcorrection.&lt;/li&gt;
&lt;li&gt;You find interesting queries that result in interesting agent behavior, and use these traces to evolve your reward function for reinforcement learning.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;next&quot;&gt;&lt;a href=&quot;#next&quot;&gt;Next&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;AI product teams use Distributional to address these problems, and many more. Distributional is a free, open, and installable platform for agent analytics. &lt;a href=&quot;https://docs.dbnl.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Try it today&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; and quickly learn how it complements your existing agent observability stack. We are also always happy to learn more about your use case and enterprise needs, so reach out to &lt;a href=&quot;mailto:nick-dbnl@distributional.com&quot;&gt;contact@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>The four stages of AI observability maturity</title><link>https://distributional.com/blog/the-four-stages-of-ai-observability-maturity</link><guid isPermaLink="true">https://distributional.com/blog/the-four-stages-of-ai-observability-maturity</guid><description>The four stages of AI observability maturity: the aggregate trap, the manual-investigation plateau, the scaling threshold, and behavioral signals. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Tue, 05 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Most AI teams think they have observability. What they actually have is instrumentation.&lt;/p&gt;
&lt;p&gt;Latency dashboards. Error rates. Cost per call. These are useful numbers — but they’re averages, and averages lie. They tell you that something changed. They don’t tell you what changed, for whom, or why.&lt;/p&gt;
&lt;p&gt;Real observability means understanding the behavioral patterns inside your production data — the user segments where your agent degrades, the query types it handles inconsistently, the failure modes that weren’t anticipated. Most teams never get there. Not because they’re not trying, but because they’re measuring the wrong things.&lt;/p&gt;
&lt;p&gt;After working with AI teams across the industry, we’ve identified four distinct stages of observability maturity. Here’s what each one looks like — and where the transitions break down.&lt;/p&gt;
&lt;h2 id=&quot;stage-1-the-aggregate-trap&quot;&gt;&lt;a href=&quot;#stage-1-the-aggregate-trap&quot;&gt;&lt;strong&gt;Stage 1: The aggregate trap&lt;/strong&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;You have monitoring. You have dashboards. When your error rate spikes, you know about it.&lt;/p&gt;
&lt;p&gt;What you don’t have: any understanding of behavioral variation.&lt;/p&gt;
&lt;p&gt;Imagine your agent starts giving worse answers to enterprise customers — but your aggregate quality score stays flat because your SMB customers are fine. You’d never see it. The signal is buried inside the average.&lt;/p&gt;
&lt;p&gt;Teams in the Aggregate Trap typically discover problems one of two ways: users report them, or a blunt threshold alert fires. Both are lagging indicators. By the time you know something is wrong, it’s been wrong for a while.&lt;/p&gt;
&lt;p&gt;The core issue isn’t that your monitoring is bad — it’s that aggregate metrics can only detect changes that affect everyone, uniformly, at the same time. Real production AI problems rarely work that way.&lt;/p&gt;
&lt;h2 id=&quot;stage-2-the-manual-investigation-plateau&quot;&gt;&lt;a href=&quot;#stage-2-the-manual-investigation-plateau&quot;&gt;&lt;strong&gt;Stage 2: The manual investigation plateau&lt;/strong&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This is where most teams with “good” observability actually live — and it’s the stage that’s hardest to recognize from the inside, because things feel like they’re working.&lt;/p&gt;
&lt;p&gt;You have tracing. You have evals. You might have custom segments you track — query categories, user cohorts, topic buckets. When something goes wrong, you can pull traces and investigate. You’re proactive about reviewing sampled data.&lt;/p&gt;
&lt;p&gt;The problem is the word sampled. And the word predefined.&lt;/p&gt;
&lt;p&gt;When you sample traces, you’re making a bet that the interesting patterns are distributed randomly across your data. They’re not. Anomalous behavior clusters — and random sampling will systematically miss low-frequency, high-impact patterns.&lt;/p&gt;
&lt;p&gt;When you use predefined segments, you can only find patterns you already suspected existed. The entire category of unknown unknowns — the behavioral clusters that would change how you prioritize your roadmap — stays invisible.&lt;/p&gt;
&lt;p&gt;You’re doing real work. You’re just doing it inside a box whose walls you can’t see.&lt;/p&gt;
&lt;h2 id=&quot;stage-3-the-scaling-threshold&quot;&gt;&lt;a href=&quot;#stage-3-the-scaling-threshold&quot;&gt;&lt;strong&gt;Stage 3: The scaling threshold&lt;/strong&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Teams at this stage are doing a lot right. Enriched trace data. Custom NLP metrics. Topic classifications. LLM-as-judge scores. Proactive behavioral investigation.&lt;/p&gt;
&lt;p&gt;The challenge is that none of it scales.&lt;/p&gt;
&lt;p&gt;Your data science team runs ad hoc analyses. They find patterns, document them, add them to dashboards. Two months later, user behavior has shifted and the patterns they documented no longer reflect what’s actually happening in production. The analysis is always chasing the product.&lt;/p&gt;
&lt;p&gt;The other scaling problem: coverage. As trace volume grows, the percentage of your data that any human actually looks at approaches zero. You can have excellent analytical frameworks and still miss the pattern that matters most this week, because it only shows up in 0.3% of traces — which is thousands of examples at scale, but invisible to any sampling strategy.&lt;/p&gt;
&lt;p&gt;The teams who break through this stage are the ones who stop asking “how do we analyze our data better?” and start asking “how do we make pattern discovery continuous and automatic?”&lt;/p&gt;
&lt;h2 id=&quot;stage-4-behavioral-signal-pioneer&quot;&gt;&lt;a href=&quot;#stage-4-behavioral-signal-pioneer&quot;&gt;&lt;strong&gt;Stage 4: Behavioral signal pioneer&lt;/strong&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Very few teams operate here. The ones that do share a few characteristics.&lt;/p&gt;
&lt;p&gt;First, pattern discovery is automated and continuous. New behavioral signals are surfaced from production data without anyone having to think to look for them. The system discovers the unknown unknowns.&lt;/p&gt;
&lt;p&gt;Second, production data drives curation. Instead of randomly sampling traces for fine-tuning or eval sets, they select based on discovered behavioral signals — which means their training data actually reflects the failure modes that matter.&lt;/p&gt;
&lt;p&gt;Third, their understanding of agent behavior compounds over time. Every week they know more about how their product behaves in production than they did the week before, in a systematic way that accumulates rather than churning.&lt;/p&gt;
&lt;p&gt;The result is that improvement cycles get faster, not slower, as the product scales.&lt;/p&gt;
&lt;h2 id=&quot;why-the-gap-between-stage-2-and-stage-4-is-so-hard-to-close&quot;&gt;&lt;a href=&quot;#why-the-gap-between-stage-2-and-stage-4-is-so-hard-to-close&quot;&gt;&lt;strong&gt;Why the gap between stage 2 and stage 4 is so hard to close&lt;/strong&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The Manual Investigation Plateau is seductive. You feel like you’re on top of your data. You have processes. You’re doing the work.&lt;/p&gt;
&lt;p&gt;What makes it hard to leave is that the problems it creates are invisible by definition. You don’t know what patterns you’re missing. You don’t see the user segments you’re not segmenting. The unknown unknowns don’t show up in your dashboards as gaps — they just don’t show up at all.&lt;/p&gt;
&lt;p&gt;Closing the gap requires a different kind of tool: one that analyzes behavioral dimensions across your entire production dataset, continuously, without requiring you to define what you’re looking for in advance.&lt;/p&gt;
&lt;h2 id=&quot;where-does-your-team-stand&quot;&gt;&lt;a href=&quot;#where-does-your-team-stand&quot;&gt;&lt;strong&gt;Where does your team stand?&lt;/strong&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;We built a 7-question self-assessment to help AI teams find out which stage they’re actually at — not which stage they think they’re at.&lt;/p&gt;
&lt;p&gt;It takes 2 minutes. Each answer reveals an insight about what your current approach is missing. At the end, you get a scored result with a specific recommendation for how to move to the next stage.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://distributional.com/assessment&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Take the assessment → distributional.com/assessment&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Enrichments versus evals</title><link>https://distributional.com/blog/enrichments-versus-evals</link><guid isPermaLink="true">https://distributional.com/blog/enrichments-versus-evals</guid><description>Evals as narrow known-known criteria versus enrichments as many weak signals for unknown-unknown discovery in production agent logs. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Tue, 28 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;I often get asked some variation of the questions, “What standard evals do you offer in your product?” Or “How does this compare to evals tooling?” My answer often starts with reframing what we do from evals to enrichments. So what do I mean by this?&lt;/p&gt;
&lt;p&gt;This post was heavily inspired by a conversation between Scott Clark, Distributional co-founder and CEO, and Jason Liu of the Developer Experience team at OpenAI. You can watch this here (especially minutes 20-35):&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=QBaSnETJr4Q&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;how-evals-fit-with-analytics&quot;&gt;&lt;a href=&quot;#how-evals-fit-with-analytics&quot;&gt;How evals fit with analytics&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The purpose of evals is to closely represent as narrowly defined criteria for success or performance as possible in an offline testing scenario. Fewer, specific, heavily scoped metrics is the objective, so you have a way to understand agent behavior on known knowns. The goal is to use these evals to optimize performance.&lt;/p&gt;
&lt;p&gt;The purpose of analytics is to find otherwise hidden signals in your production agent logs. More, general specific, scalable metrics are the objective, so you have a way to identify unknown unknown and known unknown signals. The goal is to use analytics to identify opportunities to optimize performance with evals and other techniques. Instead of trying to come up with an eval that perfectly encapsulates quality, we use a wide range of weaker signals to find the interesting correlations that suggest unknown issues or opportunities.&lt;/p&gt;
&lt;p&gt;Philosophically, the idea behind analytics is that it is impossible to identify all the potential emergent behaviors from your agent pre-production. The only way to understand all the permutations of user, usage, agent response, and all of the context within this process is to observe what happens in production.&lt;/p&gt;
&lt;p&gt;AI product teams use analytics to identify clusters, correlations, patterns, or other interesting insights on agent behavior, and then develop evals off of these insights. Whenever they create a new eval, they add that to the combination of metrics that are assessed by the analytics solution, building richer and richer data for this analysis over time.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/enrichments-versus-evals/69ea69352e50d582841715f1_image1.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;enrichments-that-power-analytics&quot;&gt;&lt;a href=&quot;#enrichments-that-power-analytics&quot;&gt;Enrichments that power analytics&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;We use the term enrichment, because the goal is to enrich your traces with attributes that represent various aspects of agent behavior – behavioral vectors to analyze.&lt;/p&gt;
&lt;p&gt;If you only have inputs and outputs, you can get some signal from analytics. If you add richer multi-level tracing data – sessions, traces, spans – you get more interesting insights from analytics. If you join these traces with downstream KPIs or session-level events, you get even better insights from analytics. And if you add your own specific evals to the collection of metrics, you get an even stronger signal from analytics. Analytics is enhanced when you can look at the full distribution of all of these behavioral vectors in a high dimensional space, and then pull out the most interesting signal.&lt;/p&gt;
&lt;p&gt;For context, Distributional offers &lt;a href=&quot;https://docs.dbnl.com/workflow/metrics&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;out of the box&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; a variety of LLM as judge metrics and traditional natural language processing metrics to get you started. But we also make it easy to design &lt;a href=&quot;https://docs.dbnl.com/workflow/metrics/llm-as-judge-metric-templates&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;new metrics&lt;/a&gt; and join traces with downstream product KPIs so you can enhance these enrichments to produce more robust analytics.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/enrichments-versus-evals/69ea6950c4884bbc1d05b263_image2.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;towards-tab-complete-analytics&quot;&gt;&lt;a href=&quot;#towards-tab-complete-analytics&quot;&gt;Towards “tab complete” analytics&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;We do these enrichments to build a bigger haystack so analytics can do the job of finding the needles in this haystack – signals to understand, fix, and improve production agents. The goal is akin to “tab complete analytics” where instead of spending a week doing data science yourself, insights are proactively presented to you and you can quickly triage – “I care about this, let’s track it. I don’t know what this is, I’ll quickly investigate. I don’t care about this, I’ll ignore it.” The goal is to continually find and fix new issues over time to boost agent performance.&lt;/p&gt;
&lt;p&gt;We facilitate these “tab complete” analytics by performing unsupervised learning and clustering on these enrichments at scale, pulling out sub-pockets of these behavioral vector distributions, and prioritizing insights that showcase situations where something didn’t happen very frequently or where a behavior is correlated with a cost, speed, or quality issue that requires attention. These clusters are different than anomalies, because sub-peaks within a multimodal distribution are also potentially relevant, not just the peaks. And agents are incredibly multimodal. We feed these clusters into an LLM – it can be on the smaller end, some of the best we’ve seen are 70b parameter LLMs – to summarize the issue and recommend fixes.&lt;/p&gt;
&lt;p&gt;These analytics are relatively cheap to compute, and the goal is to engage with them to refine the enrichments and analysis over time. As you improve and fix your agent, these insights will adapt to find new signals that need your attention, and around you’ll go on this flywheel.&lt;/p&gt;
&lt;h2 id=&quot;next&quot;&gt;&lt;a href=&quot;#next&quot;&gt;Next&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Before founding Distributional, we spent years together focused on Bayesian optimization for hyperparameter optimization and other offline experimentation tasks when training traditional AI/ML models at a startup called &lt;a href=&quot;https://sigopt.org/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;SigOpt&lt;/a&gt; that we sold to Intel in 2020. We worked with American Express, Netflix, OpenAI, $1T worth of hedge funds, and many others operating at the extreme end of scale and performance for these systems. In all cases, when they started optimizing, they’d find they had new issues they had to account for in their objective function. The way we’ve designed this version of analytics for production agents is to enable you to continuously refine this objective function – a continuous behavioral space – so you can properly optimize agents across all attributes and guide their behavior over time. You optimize what you measure, but you only measure what you know.&lt;/p&gt;
&lt;p&gt;Distributional is a free, open, and installable platform for agent analytics. &lt;a href=&quot;https://docs.dbnl.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Try it today&lt;/a&gt; and quickly learn how it complements your existing agent observability stack. We are also always happy to learn more about your use case and enterprise needs, so reach out to &lt;a href=&quot;mailto:nick-dbnl@distributional.com&quot;&gt;contact@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>The agent observability hierarchy</title><link>https://distributional.com/blog/the-agent-observability-hierarchy</link><guid isPermaLink="true">https://distributional.com/blog/the-agent-observability-hierarchy</guid><description>The agent observability hierarchy: logging, monitoring, analytics; why scale and agent complexity broke trace-reading, and how the layers complement. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Tue, 21 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;With seemingly exponential acceleration of foundation model capabilities, it is becoming easier for AI product teams to build performant agents. These teams are now being tasked with managing these agents at a greater scale and a wider variety of use cases than previously contemplated. To do this well, they need to implement an &lt;a href=&quot;https://www.nvidia.com/en-us/glossary/data-flywheel/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;AI data flywheel&lt;/a&gt; powered by a modern, agent-first approach to observability.&lt;/p&gt;
&lt;p&gt;This post was heavily inspired by a conversation between Scott Clark, Distributional co-founder and CEO, and Jason Liu of the Developer Experience team at OpenAI. You can watch this here (especially minutes 1 - 15):&lt;/p&gt;
&lt;h2 id=&quot;agent-observability-hierarchy&quot;&gt;&lt;a href=&quot;#agent-observability-hierarchy&quot;&gt;Agent observability hierarchy&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Scale alone breaks the former paradigm of “I can look at most of my traces to understand how my agent is behaving” that used to be the gold standard of online agent observability. But to complicate things further, to leverage these enhanced capabilities of the foundation models it makes sense to let them make many more decisions about the workflow to get to an answer—which creates a mess of tool calls, context retrieval, multi-prompt chains, and multiple LLM calls with no human in the loop. Add agent complexity with this scale, and agent observability is fully broken.&lt;/p&gt;
&lt;p&gt;These AI product teams need a modern agent observability stack of capabilities. At Distributional, we think of this as a three-component hierarchy of agent observability: Logging, Monitoring, and Analytics.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-agent-observability-hierarchy/69d42d1c718e95a19fc2b85f_image3.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;how-analytics-fits-with-logging-and-monitoring&quot;&gt;&lt;a href=&quot;#how-analytics-fits-with-logging-and-monitoring&quot;&gt;How analytics fits with logging and monitoring&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Most teams have already started to repair their agent observability stack. Richer agent logging that includes sessions, traces, and spans is widely available through OpenTelemetry and plenty of open source products that support it. Teams often start logging to build and debug their agents. These logs are fed into an eval tool that makes it easy to iterate on prompts or design custom functions to score or classify outputs. Logging what has happened is the first step of any observability stack, and for agents it is critical to have multiple levels reported (session, trace, and span).&lt;/p&gt;
&lt;p&gt;Similarly, AI product teams have invested in agent-specific monitoring that tracks, thresholds, and gives fast response alerts on deviations from thresholds on known cost, speed, and quality metrics. These monitoring solutions are designed to trade off richness of analysis for timeliness—you need to know when cost has spiked or your agent is down. These have been redesigned to account for LLM-as-judge metrics and agent specific evals, as well as tokens, feedback, and other standard speed, cost, and quality metrics for agents.&lt;/p&gt;
&lt;p&gt;But there is a third component that every team needs to implement to complete the AI feedback lifecycle: &lt;strong&gt;Analytics&lt;/strong&gt;. Think of analytics as similar to product analytics for a traditional software product, but instead of treating the &lt;em&gt;user&lt;/em&gt; as the atomic unit you treat the &lt;em&gt;agent&lt;/em&gt; as such. Whereas monitoring trades off richness for timeliness, analytics trades off timeliness for richness.&lt;/p&gt;
&lt;p&gt;Analytics are particularly important for agents due to their complexity. Logging and monitoring is sufficient if you know everything you need to track for a relatively straightforward software product. But when you are deferring cognitive tasks to the agent and the agent is deciding which tools, context, models, or steps to take, it is impossible to know all the issues that could arise. Similarly, most agents are flexible enough that users could make a wide variety of diverse requests. Without knowing in advance how users are going to use the agent or how the agent will behave when they do, there are plenty of unknown unknowns that could have a significant impact on agent performance.&lt;/p&gt;
&lt;p&gt;If you pair logging and monitoring with analytics, you can complete this feedback loop. You log traces to provide information on agent behavior, monitoring to threshold known issues, and analytics to discover new issues to be fixed and monitored.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-agent-observability-hierarchy/69d42d3ece457126dfd1dcbe_image1.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;benefits-of-this-approach-to-agent-observability&quot;&gt;&lt;a href=&quot;#benefits-of-this-approach-to-agent-observability&quot;&gt;Benefits of this approach to agent observability&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Agent-specific logging and monitoring are a good start, and can begin to give you visibility into whether things are working. By adding analytics to the agent observability stack, however, you enable continuous improvement of the agent. Analytics surfaces signals that are otherwise unknown, and supports these signals with contextual evidence—specific logs that represent the issue or opportunity to improve.&lt;/p&gt;
&lt;p&gt;Armed with this evidence, AI product teams can add these specific logs to their offline evals, reward functions for reinforcement learning pipelines, or datasets for fine-tuning processes to optimize agent performance. These analytics can also guide simpler fixes, such as prompt improvements, context refinement, or tool adjustments that can boost performance.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-agent-observability-hierarchy/69d42d561c9042d10f6747a6_image2.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;next&quot;&gt;&lt;a href=&quot;#next&quot;&gt;Next&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional is a free, open, and installable platform for agent analytics. &lt;a href=&quot;https://docs.dbnl.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Try it today&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; and quickly learn how it complements your existing agent observability stack. We are also always happy to learn more about your use case and enterprise needs, so reach out to &lt;a href=&quot;mailto:nick-dbnl@distributional.com&quot;&gt;contact@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Demo: Financial analyst</title><link>https://distributional.com/blog/demo-financial-analyst</link><guid isPermaLink="true">https://distributional.com/blog/demo-financial-analyst</guid><description>A financial research agent under behavioral analytics: baseline understanding, proactive insights, and active governance for regulated industries. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Fri, 17 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;In a recent conversation with an innovation leader at one of the largest asset managers, he put his problem quite simply, “How can I get my team to move from passive monitoring to active governance of their agents in production?” His view was that analytics on production traces could complement their existing reactive monitoring stack to give them a more proactive way to probe and assess what needed to be fixed, could be improved, or was even happening with their agent in the first place.&lt;/p&gt;
&lt;p&gt;His experience inspired our demo with a much more streamlined financial research agent. In this example, we’ll show you how to use Distributional’s (DBNL’s) unsupervised, automated analytics on production agent traces to drive more active governance of an agent in production.&lt;/p&gt;
&lt;p&gt;To explore this data using our read only SaaS demo environment, go to &lt;a href=&quot;https://app.dbnl.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://app.dbnl.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;, use these credentials, and select the project “Financial Research Agent Example”:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Username: &lt;code&gt;demo-user@distributional.com&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Password: &lt;code&gt;dbnldemo1!&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And to test the product yourself, you can &lt;a href=&quot;https://docs.dbnl.com/get-started/quickstart#deploy-a-local-sandbox-with-example-data&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;install our sandbox to get started&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;
&lt;h2 id=&quot;setup&quot;&gt;&lt;a href=&quot;#setup&quot;&gt;Setup&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;We built a relatively streamlined financial research agent that uses gpt-4o-mini to call a variety of tools (stock symbol, stock price, stock news, company info) to answer questions related to a company’s financial performance. We simulated a few hundred queries per day to provide a reasonable distribution. I grabbed the first session in the logs that Distributional analyzed to provide a sense of these queries.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=_r26nxRbwHM&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;understanding-baseline-behavior-of-this-agent&quot;&gt;&lt;a href=&quot;#understanding-baseline-behavior-of-this-agent&quot;&gt;Understanding baseline behavior of this agent&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Out of the box, Distributional provides a variety of useful information to help understand agent behavior in the context of usage patterns. Distributional summarizes standard cost, speed, and quality metrics that you’d see in any passive monitoring system to ground other analytics in this context. But DBNL builds on this foundation in a few ways.&lt;/p&gt;
&lt;p&gt;After seven days of data, DBNL produces topics and classifies requests into these topic categories. These topics represent user intent, and DBNL shows how the propensity of these topics shift over time. In this case, most users are asking for detailed profiles of a single company, recent news that impacted a company’s stock, or comparisons for multiple companies across financial and operational metrics.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42b65ffd7eb45dcbaff88_image4.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Along with topics, DBNL also automatically produces a tool call sequence flow. This flow gives a summary of how the agent is calling tools, and can be filtered by interesting segments of logs.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42b9b3074a2f2360c9e7c_image2.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;DBNL also grounds topics, which offer insight on inputs, in the context of downstream metrics and KPIs that offer insight on outputs. This provides a quick sense of whether there is a potential issue with a given type of query. In this case, DBNL assesses average feedback score by topic to provide this view of input-output consistency.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42baf47e72c51ef7d25da_image7.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Distributional also produces a variety of standard LLM-as-judge metrics out of the box that track response quality across all production agent traces. DBNL sets you up to use your own model for these metrics, but also includes recommendations in the docs for models that work particularly well and efficiently for these analytics.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42bc5ecab11e6809a1cd7_image5.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;proactive-insights-on-agent-behavior&quot;&gt;&lt;a href=&quot;#proactive-insights-on-agent-behavior&quot;&gt;Proactive insights on agent behavior&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;DBNL provides this deeper understanding of agent behavior to ground more proactive insights on how this behavior is evolving over time, and what actions your team should take to fix or improve the agent. Each day, DBNL produces a new collection of Insights with a quick summary of the type of insight, the KPI it impacts, and its severity for rapid triage.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42bd90cad280f15261eb8_image9.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;If we click into the “Incorrect or Incomplete Data Retrieval” insight, we get a deeper summary and selection of example logs that provide evidence supporting the insight. You see this insight refers to issues related to pulling the right stock symbol and other types of data retrieval issues with the agent.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42bedc39ed3c4ff45194c_image8.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Embedded in this same insight card is also a set of recommendations with a light estimate of level of effort. This example shows prompt, orchestration, or tool changes that could improve performance.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42c0913c695e3542e6e6d_image1.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Finally, the insight card also recommends a segment you can create to track the issue. This is particularly helpful for confirming that a change is having the desired impact—you’ll be able to see the relevant metrics for this particular segment respond to the change, or know that you need to try something else instead.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42c2a42b9b0083e2c7898_image3.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;It is worth noting that you can also derive your own insights intuitively with DBNL by filtering the logs from the Dashboard or Explorer pages. Here is a quick example of filtering the financial research agent logs by output = irrelevant.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42c449b682538aec38784_image10.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;When you dig in, you quickly see many examples where the agent is reporting it couldn’t find the stock symbol, which may suggest an issue with the tools or how these tools are being called by the agent.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-financial-analyst/69d42c595d798b72c717c911_image6.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;get-started&quot;&gt;&lt;a href=&quot;#get-started&quot;&gt;Get started&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The goal of using DBNL’s analytics on production agent traces is to help you find the unknown unknowns so you can proactively fix, improve, or change your agent. These daily insights on agent behavior complement logging and monitoring to complete the AI observability stack. And these analytics facilitate a more active approach to agent governance, which, in turn, gives AI product teams the tools they need to confidently scale agent usage in production.&lt;/p&gt;
&lt;p&gt;The easiest way to &lt;a href=&quot;https://docs.dbnl.com/get-started/quickstart&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;get started&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; is to use a free SaaS demo account to review this example and other examples that we’ve pre-loaded in Distributional. Next, you can install our &lt;a href=&quot;https://docs.dbnl.com/platform/deployment/sandbox&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;sandbox&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; locally on your laptop in ten minutes and run through a tutorial that shows you how to use Distributional for a toy example. Once more familiar with our functionality, you can &lt;a href=&quot;https://docs.dbnl.com/platform/deployment&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;install&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; the full service for free using a Terraform Module or Helm Chart. We are happy to help through any step of this process, so reach out at &lt;a href=&quot;mailto:support@distributional.com&quot;&gt;support@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Demo: Improving agents with production data analysis</title><link>https://distributional.com/blog/demo-improving-agents-with-production-data-analysis</link><guid isPermaLink="true">https://distributional.com/blog/demo-improving-agents-with-production-data-analysis</guid><description>The open NVIDIA stack (NeMo Agent Toolkit, NIM, Optimizer) on a production agent: finding improvement signals in production data. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 08 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;This demo walks through how to use &lt;a href=&quot;http://docs.dbnl.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Distributional&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with &lt;a href=&quot;https://docs.nvidia.com/nim-operator/latest/index.html&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NVIDIA NIMs&lt;/a&gt; and &lt;a href=&quot;https://docs.nvidia.com/nemo/agent-toolkit/latest/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NVIDIA NeMo Agent Toolkit&lt;/a&gt; as a completely open stack to scale analytics for a production outing agent. This agent is designed to reply to questions on outings with a full itinerary for the user, including links for maps and reservations. In the process of developing this itinerary, the agent calls a variety of tools.&lt;/p&gt;
&lt;p&gt;If you prefer watching a video instead, we’ve walked through this example in a &lt;a href=&quot;https://distributional.com/blog/demo-understand-ai-agents-with-better-analytics&quot;&gt;demo video&lt;/a&gt; &lt;em&gt;[NOTE: That post was not restored in the selected archive.]&lt;/em&gt;. This example runs off of the same outing agent that is used for our quickstart, which we’ve also &lt;a href=&quot;https://distributional.com/blog/discovering-and-fixing-issues-with-an-ai-agent&quot;&gt;posted about in the past&lt;/a&gt; &lt;em&gt;[NOTE: That post was not restored in the selected archive.]&lt;/em&gt; if you are looking for a video explanation instead.&lt;/p&gt;
&lt;p&gt;To explore this data using our read only SaaS demo environment, go to &lt;a href=&quot;https://app.dbnl.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://app.dbnl.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;, use these credentials, and select the project “Outing Agent: NeMo Agent Toolkit AWS re:Invent Example”:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Username: &lt;code&gt;demo-user@distributional.com&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Password: &lt;code&gt;dbnldemo1!&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;setup&quot;&gt;&lt;a href=&quot;#setup&quot;&gt;Setup&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This example relies on an open, free stack that includes NVIDIA and Distributional components. Distributional provides the analytics layer, and is designed to integrate seamlessly regardless of agent framework, model, or optimizer your team uses.&lt;/p&gt;
&lt;p&gt;In this example, we rely on the NVIDIA &lt;a href=&quot;https://docs.nvidia.com/nemo/agent-toolkit/latest/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NeMo Agent Toolkit v1.3.0&lt;/a&gt; as our extensible framework, &lt;a href=&quot;https://docs.nvidia.com/nim-operator/latest/index.html&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NVIDIA NIM k8s Operator&lt;/a&gt; to efficiently run our LLMs for both the agent and analytics tasks, and &lt;a href=&quot;https://docs.nvidia.com/nemo/agent-toolkit/latest/reference/optimizer.html&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NVIDIA NeMo Agent Toolkit Optimizer&lt;/a&gt; for offline hyperparameter optimization to improve the agent guided by insights from Distributional. We used AWS S3 for storing traces from the agent, and AWS EC2 for compute.&lt;/p&gt;
&lt;p&gt;As it runs, NAT publishes 7k traces per day, which DBNL analyzes. DBNL enriches these traces with LLM-as-judge and standard metrics, then analyzes these metrics to uncover behavioral signals hidden in these logs. DBNL uses gpt-oss-20B for LLM-as-judge, and efficiently scales this judge by using a NIM on a p5 instance. As DBNL discovers new signals, we use functionality in NAT like their HPO feature to make improvements or fixes to the agent guided by insights from these signals.&lt;/p&gt;
&lt;h2 id=&quot;find-signals-for-improving-production-agents&quot;&gt;&lt;a href=&quot;#find-signals-for-improving-production-agents&quot;&gt;Find signals for improving production agents&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional provides the same summary metrics you’d see in an LLM monitoring dashboard on speed, cost, and quality of the agent. But the power in Distributional is augmenting these metrics with downstream product KPIs (e.g., revenue, user feedback, click rate, etc.), richer metrics on agent behavior (e.g., tool calls, tool sequence flows, etc.), and usage patterns (e.g., topics, intent, etc.). As Distributional automatically clusters and correlates across this richer panel of metrics, it produces both a richer understanding of agent behavior, and insights on how this behavior evolves over time.&lt;/p&gt;
&lt;p&gt;As an example of richer understanding of baseline agent behavior from this example, you can see that crossing user feedback (a downstream product KPI) with topics (usage pattern) you can get a quick sense that there are issues in how the agent is performing, but it isn’t related to the user query – it seems to extend across a variety of topics that each have an average feedback score of 1 (and others with scores in the 3s, which may also be troublesome).&lt;/p&gt;
&lt;p&gt;Similarly, metrics like output relevancy and user frustration show a clear picture that there is a relatively high propensity of irrelevancy and frustration in the first few days running the agent.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-improving-agents-with-production-data-analysis/69d4296582382fb0d145fb61_image4.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Distributional complements this richer baseline understanding of agent behavior with daily, targeted, actionable insights useful to guide prioritization of work to fix or improve the agent on an ongoing basis. In this case, on the second data running, Distributional picks up on an issue with link correctness from the agent that may be a function of the system prompt or how the agent is calling its tools to produce the itinerary output.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-improving-agents-with-production-data-analysis/69d429840baca3885186eb3e_image5.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Distributional provides a link to view these filtered logs so you can directly review examples of this issue. Pretty quickly, you see examples where there should be links for making a reservation, for example, that are missing. There are also examples where there is no link, or there is an incorrect link.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-improving-agents-with-production-data-analysis/69d42997c57ff17b60a25306_image6.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Armed with these insights, we pull them offline to expand the eval set, and then run automatic selection of a new prompt using the &lt;a href=&quot;https://docs.nvidia.com/nemo/agent-toolkit/latest/reference/optimizer.html&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NVIDIA NeMo Agent Toolkit Optimizer&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;track-the-improvement-over-time&quot;&gt;&lt;a href=&quot;#track-the-improvement-over-time&quot;&gt;Track the improvement over time&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;In some cases, you may want to run a classifier for a few days to confirm the issue rather than fixing it immediately. Distributional provides templated metrics that require just a few clicks to customize. In this case, we create a custom classifier for link correctness.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-improving-agents-with-production-data-analysis/69d429b28da5de7f90e91f59_image3.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;This metric is also valuable for confirming an improvement after a change has been made. In the Metrics tab of the DBNL Dashboard, you see the error rate for link correctness drop from 34% to 14% after we’ve selected the new prompt.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-improving-agents-with-production-data-analysis/69d429c6d9df84debfb9b4ba_image1.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Even if you don’t create a custom metric, Distributional will show temporal change in the variety of behavioral metrics that should reflect this improvement as well. In this case, we see average feedback score jump from just over 2 to 4.5 after this change.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-improving-agents-with-production-data-analysis/69d429dceaa9f4c9410f39aa_image2.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;We also see output irrelevancy fall from 13% to 6% and average user frustration dip from 1.7 to 1.0 after the change.&lt;/p&gt;
&lt;h2 id=&quot;get-started&quot;&gt;&lt;a href=&quot;#get-started&quot;&gt;Get started&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This example shows how to take Distributional’s insights from analyzing production agent traces, and use them in an offline optimization workflow to tweak a prompt for immediate performance gains. Whether it is optimizing a prompt, hyperparameter, context, tools, or models with reinforcement learning, fine tuning, or re-training, Distributional can make this process a production data-driven feedback cycle.&lt;/p&gt;
&lt;p&gt;The easiest way to &lt;a href=&quot;https://docs.dbnl.com/get-started/quickstart&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;get started&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; is to use a free SaaS demo account to review this example and other examples that we’ve pre-loaded in Distributional. Next, you can install our &lt;a href=&quot;https://docs.dbnl.com/platform/deployment/sandbox&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;sandbox&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; locally on your laptop in ten minutes and run through a tutorial that shows you how to use Distributional for a toy example. Once more familiar with our functionality, you can &lt;a href=&quot;https://docs.dbnl.com/platform/deployment&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;install&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; the full service for free using a Terraform Module or Helm Chart. We are happy to help through any step of this process, so reach out at &lt;a href=&quot;mailto:support@distributional.com&quot;&gt;support@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Demo: Agent hyperparameter optimization with behavioral analytics</title><link>https://distributional.com/blog/demo-agent-hyperparameter-optimization-with-behavioral-analytics</link><guid isPermaLink="true">https://distributional.com/blog/demo-agent-hyperparameter-optimization-with-behavioral-analytics</guid><description>Closing the loop from analytics to optimization: diagnosing a seeded bug and driving hyperparameter optimization from behavioral signals. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Mon, 06 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Sometimes, simpler is better. In this demo, we simplify the example to show you more concisely how our product works to provide behavioral analytics that facilitate a hyperparameter optimization loop to boost performance of an agent in production.&lt;/p&gt;
&lt;p&gt;To run this example locally using your own sandbox instance of DBNL, go to: &lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/nemo_agent_toolkit_hpo_example&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://github.com/dbnlAI/examples/tree/main/nemo_agent_toolkit_hpo_example&lt;/a&gt;. In this example, there is also a &lt;a href=&quot;https://camo.githubusercontent.com/cd0231483081dbf761a6b36ef0b9a8c879b6abb34c9e64e19e97df1151c7c447/68747470733a2f2f636f6e74656e742e676974626f6f6b2e636f6d2f636f6e74656e742f65784d367655304448644c4837547952677648392f626c6f62732f34456c777347525377706e72583141434c6d586a2f68706f5f64656d6f5f736d616c6c5f6f70742e676966&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;video walkthrough&lt;/a&gt; of the example as well.&lt;/p&gt;
&lt;p&gt;To explore this data using our read only SaaS demo environment, go to &lt;a href=&quot;https://app.dbnl.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://app.dbnl.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;, use these credentials, and select the project “Google ADK Calculator Hyperparameter Optimization Example”:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Username: &lt;code&gt;demo-user@distributional.com&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Password: &lt;code&gt;dbnldemo1!&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-agent-hyperparameter-optimization-with-behavioral-analytics/698d0df28bc2cf70065c2ba7_HPO_Demo.gif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;setup&quot;&gt;&lt;a href=&quot;#setup&quot;&gt;Setup&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;We built a simple calculator agent built using &lt;a href=&quot;https://google.github.io/adk-docs/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Google Agent Development Kit&lt;/a&gt; (ADK) that calls a variety of tools (multiplication, addition, subtraction, etc) to answer math questions. This example isn’t intended to mimic what you will be doing in the wild, but to give you a quick sense of the power of Distributional in the production agent stack to complete a positive feedback cycle.&lt;/p&gt;
&lt;p&gt;We wrapped this &lt;a href=&quot;https://google.github.io/adk-docs/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Google ADK&lt;/a&gt; agent with the &lt;a href=&quot;https://github.com/NVIDIA/NeMo-Agent-Toolkit&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NVIDIA NeMo Agent Toolkit&lt;/a&gt;, both available openly to anyone developing agents. We also intentionally added an issue to the agent, in this case a sigmoid function that will cause errors that increase in severity as larger numbers are added or multiplied.&lt;/p&gt;
&lt;h2 id=&quot;diagnosing-the-issue&quot;&gt;&lt;a href=&quot;#diagnosing-the-issue&quot;&gt;Diagnosing the issue&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional automatically provides high level understanding on a variety of speed, quality, and cost metrics that you may also see in an LLM monitoring tool. But Distributional combines these metrics with deeper insight on tool usage, topics, and product KPIs to provide a clearer picture of &lt;em&gt;agent behavior&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=UizipyDx0kY&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;From these analytics, Distributional makes it easy to filter the relevant logs so you can confirm any aspect of this behavior with direct evidence, and then use these data samples for offline prompt iteration, reinforcement learning, fine tuning, tool changes, or, in this case, hyperparameter optimization. Here we’ve selected all logs that the LLM-as-judge has labeled to have irrelevant output so we can get a quick picture of what is wrong.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-agent-hyperparameter-optimization-with-behavioral-analytics/698d0cec5878d81ca319d1e1_422bd387.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Distributional also automatically clusters and correlates these attributes as part of behavioral analysis to produce daily actionable insights—fixes or improvements the team should prioritize. In this case, one of the first insights relates to output imprecision, and makes recommendations for how to address the issue.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-agent-hyperparameter-optimization-with-behavioral-analytics/698d0cf8ede5aa07cb46e6e6_c3f692f7.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;We talked earlier about directly exploring filtered logs. You can also compare segments in our Explorer page to get a visual sense of what is happening. In this case, by comparing logs with absolute error and output expected &amp;gt;= 100 (which we got from the Insight), we can quickly see large numbers must be an issue.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-agent-hyperparameter-optimization-with-behavioral-analytics/698d0d0398c7ecdcf9ea7b61_38a6a653.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;data-driven-hyperparameter-optimization&quot;&gt;&lt;a href=&quot;#data-driven-hyperparameter-optimization&quot;&gt;Data-driven hyperparameter optimization&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;In this demo, we used an intentionally simplified version of hyperparameter optimization to give a clearer, more concise sense of how you’d use Distributional to run a data-driven version of this process.&lt;/p&gt;
&lt;p&gt;In this case, we use Distributional to sample production data that can then be used to provide a more robust dataset for offline hyperparameter optimization. We create a new config, run a broader dataset sampled using Distributional’s analysis, and re-ran the simple hyperparameter optimization loop to select a new parameter value (in this case using Optuna’s Tree Parzen Estimator). This job is orchestrated by &lt;a href=&quot;https://docs.nvidia.com/nemo/agent-toolkit/latest/improve-workflows/optimizer.html&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NVIDIA’s NeMo Agent Toolkit Optimizer&lt;/a&gt; functionality.&lt;/p&gt;
&lt;p&gt;Here are the steps you can run directly from this example to get perspective on how this works offline relative to and guided by Distributional’s online analysis.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-agent-hyperparameter-optimization-with-behavioral-analytics/698d0d18c619b710fbed8d72_6ec37501.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;verifying-the-hpo-driven-performance-gain&quot;&gt;&lt;a href=&quot;#verifying-the-hpo-driven-performance-gain&quot;&gt;Verifying the HPO-driven performance gain&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional is an automatic, clean, and visual way to verify that these offline changes are reflected in production. In this case we quickly see average feedback jump up from less than 2 to a steady 5, because the error has been fixed with this change. We also see output irrelevancy go to roughly zero.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-agent-hyperparameter-optimization-with-behavioral-analytics/698d0d2ebda585fd315d9922_0b97add7.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;With a single click, we also created segments to track to ensure that the issue didn’t crop back up. You see the math error segment plummets from 80% to 0% after the change and stays there.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-agent-hyperparameter-optimization-with-behavioral-analytics/698d0d3970c183c7eb3cb8d2_3c35d237.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;This was a relatively simple example designed to give you quick intuition on how to use Distributional. But this data-driven approach to leveraging Distributional analysis of production logs to guide further offline development and optimization of an agent applies to agentic products of any complexity.&lt;/p&gt;
&lt;h2 id=&quot;get-started&quot;&gt;&lt;a href=&quot;#get-started&quot;&gt;Get started&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The easiest way to &lt;a href=&quot;https://docs.dbnl.com/get-started/quickstart&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;get started&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; is to use a free SaaS demo account to review this example and other examples that we’ve pre-loaded in Distributional. Next, you can install our &lt;a href=&quot;https://docs.dbnl.com/platform/deployment/sandbox&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;sandbox&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; locally on your laptop in ten minutes and run through a tutorial that shows you how to use Distributional for a toy example. Once more familiar with our functionality, you can &lt;a href=&quot;https://docs.dbnl.com/platform/deployment&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;install&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; the full service for free using a Terraform Module or Helm Chart. We are happy to help through any step of this process, so reach out at &lt;a href=&quot;mailto:support@distributional.com&quot;&gt;support@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Demo: Fixing agent issues in production with tab complete analytics</title><link>https://distributional.com/blog/demo-fixing-agent-issues-in-production-with-tab-complete-analytics</link><guid isPermaLink="true">https://distributional.com/blog/demo-fixing-agent-issues-in-production-with-tab-complete-analytics</guid><description>Google ADK calculator tutorial: multiple ingestion pathways, circumstance-and-pathway analysis, and insight-driven fixes. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Mon, 30 Mar 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Sometimes, simpler is better. In this demo, we simplify the example to show you more concisely how our product works.&lt;/p&gt;
&lt;p&gt;We built a simple calculator agent built using &lt;a href=&quot;https://google.github.io/adk-docs/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Google Agent Development Kit&lt;/a&gt; (ADK) that calls a variety of tools (multiplication, addition, subtraction, etc) to answer math questions. We used a variety of data ingestion methods with this simple agent to show how you can use Distributional across diverse logging and tracing setups. And we introduced a couple simple issues to see if our analytics would automatically surface them. This example isn’t intended to mimic what you will be doing in the wild, but to give you a quick sense of the power of Distributional in the production agent stack.&lt;/p&gt;
&lt;p&gt;To run this example locally using your own sandbox instance of DBNL, go to our Github tutorial: &lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/adk_calculator_tutorial&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://github.com/dbnlAI/examples/tree/main/adk_calculator_tutorial&lt;/a&gt;. In this tutorial, there is also a &lt;a href=&quot;https://camo.githubusercontent.com/5b899f33b3d5e698b7ba9706c433babf3ba72a11cfc1faa72eddeb6020691ca6/68747470733a2f2f636f6e74656e742e676974626f6f6b2e636f6d2f636f6e74656e742f65784d367655304448644c4837547952677648392f626c6f62732f444d3331595235497237374679685447564a73412f63616c635f64656d6f5f736d616c6c5f6f70742e676966&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;video walkthrough&lt;/a&gt; of the example if you’d rather follow along this way.&lt;/p&gt;
&lt;p&gt;To explore this data using our read only SaaS demo environment, go to &lt;a href=&quot;https://app.dbnl.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://app.dbnl.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;, use these credentials, and select the project “Google ADK Calculator Agent Tutorial”:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Username: &lt;code&gt;demo-user@distributional.com&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Password: &lt;code&gt;dbnldemo1!&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-fixing-agent-issues-in-production-with-tab-complete-analytics/698d0da39ccee54d53ad7eb4_Demo-fixing-calculator.gif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;getting-data-into-dbnl&quot;&gt;&lt;a href=&quot;#getting-data-into-dbnl&quot;&gt;Getting data into DBNL&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;One advantage of this simpler agent demo example is that we can also quickly showcase various ways to get data into Distributional. We’ve designed Distributional to be flexible, running almost like a plug-and-play compute engine regardless of how you collect or where you store your traces. For this example, we show multiple pathways you can use to get data into DBNL depending on your circumstances:&lt;/p&gt;
&lt;h4 id=&quot;circumstance-and-pathway&quot;&gt;&lt;a href=&quot;#circumstance-and-pathway&quot;&gt;&lt;strong&gt;Circumstance and pathway&lt;/strong&gt;&lt;/a&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li&gt;Getting started:&lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/adk_calculator_sdk_from_otel&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Agent=&amp;gt;OTEL=&amp;gt;Traces JSONL=&amp;gt;Augment=&amp;gt;SDK=&amp;gt;DBNL&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Ongoing production&lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/adk_calculator_otel_direct&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;: Agent=&amp;gt;OTEL=&amp;gt;DBNL&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Existing ETL pipeline:&lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/adk_calculator_sdk_from_json&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Agent=&amp;gt;OTEL=&amp;gt;Raw JSONL=&amp;gt;SDK=&amp;gt;DBNL&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Traces stored already: &lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/sdk_from_langfuse_export&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Agent=&amp;gt;OTEL=&amp;gt;Langfuse=&amp;gt;Export File=&amp;gt;Augment=&amp;gt;SDK=&amp;gt;DBNL&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;understanding-agent-behavior&quot;&gt;&lt;a href=&quot;#understanding-agent-behavior&quot;&gt;Understanding agent behavior&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;We recommend you start by uploading a week’s worth of full session, trace, and span data to Distributional to get the richest insights. By analyzing this data, Distributional first helps you &lt;a href=&quot;https://docs.dbnl.com/workflow/dashboards&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;understand the status&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; of the behavior of your users, your agent, and your outcomes in production.&lt;/p&gt;
&lt;p&gt;This analysis includes a greater variety of cost, quality, and speed metrics you may find in LLM monitoring tools, including agent specific metrics around tool calls, tool error rates, tool sequences, user frustration, and response quality. But it also combines these metrics with product outcome metrics via downstream joins of these KPIs with your trace data. Further, it combines these metrics and KPIs with insights on user behavior, classifying their queries by topic. Finally, the analysis includes clustering and correlating across these metrics, KPIs, and usage patterns to give you a more complete picture of overall agent behavior.&lt;/p&gt;
&lt;p&gt;Here is an example of some of that richer information Distributional gives you on user intent and agent behavior from these calculator tool calls.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-fixing-agent-issues-in-production-with-tab-complete-analytics/698d0b25294fc9c8519cbab8_2b8e94c1.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;And here is an example where we bring in a downstream product KPI like user feedback and correlate it with topics to understand which topics are performing better or worse according to user feedback. We get a quick view that adding or multiplying two numbers may be an issue for this calculator.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-fixing-agent-issues-in-production-with-tab-complete-analytics/698d0b30977dde9915079568_8a08391e.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Distributional also has an &lt;a href=&quot;https://docs.dbnl.com/workflow/explorer&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Explorer&lt;/a&gt; page that can be used to analyze single segments, compare multiple segments, or look at segment comparisons temporally. This page is useful for building intuition on metrics, segments, and behavior when you are getting started. And can offer an easy way to analyze a/b tests once you are further along, ensuring that any changes you make to your agents are performing as expected.&lt;/p&gt;
&lt;p&gt;As you learn from the Explorer page, Distributional also makes it easy to create new &lt;a href=&quot;https://docs.dbnl.com/workflow/metrics&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;metrics&lt;/a&gt; off of judge templates, or define new &lt;a href=&quot;https://docs.dbnl.com/workflow/segments&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;segments&lt;/a&gt; to track that reflect interesting clusters of usage or agent behavior. In this example, we’ve quickly added segments for various math operations so we can track error rates across them cleanly and get a quick visual on what is and isn’t working.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-fixing-agent-issues-in-production-with-tab-complete-analytics/698d0b406e04e0a6ee9e1126_f741fa13.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;fixing-issues-with-insights-and-recommendations&quot;&gt;&lt;a href=&quot;#fixing-issues-with-insights-and-recommendations&quot;&gt;Fixing issues with insights and recommendations&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Building on core understanding around agent behavior, one of the primary benefits of Distributional is that it surfaces new insights each day on the performance, behavior, cost, quality, speed, or usage of your agent, as well as the correlations across these attributes. This gives your AI product team a fresh set of potential issues or opportunities for improvement that can guide additional work on the agent to ensure you maintain high performance in production. And all of this happens automatically without taking valuable time from your team.&lt;/p&gt;
&lt;p&gt;In this calculator example, we’ve intentionally introduced two issues so you can see how the product works:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Trying to add any number larger than 90 will result in a tool error&lt;/li&gt;
&lt;li&gt;Multiplying numbers where the first number is less than 10 will cause an arithmetic miscalculation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Distributional Insights are designed to give you a toehold into understanding these issues, and a starting point for how to resolve them with tailored recommendations.&lt;/p&gt;
&lt;p&gt;Here is an example of a single day of insights from running this calculator. Distributional has generated each of these insights automatically, and classified them by KPI (quality, speed, cost), type (error, issue, change) and severity (high, medium, low).&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-fixing-agent-issues-in-production-with-tab-complete-analytics/698d0b6717f71cbb83c4fc3c_0c6071d4.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;This makes triage easy, as you typically will want to start with the high severity issue, which is also the first insight listed. You can immediately get a sense of what the issue is from the summary. Then by clicking in, you get more rich information, including a description of the issue, example logs to explore as evidence, and recommendations for how to fix it. We also offer single-click segment creation to easily track this issue over time.&lt;/p&gt;
&lt;p&gt;In this case, the calculator is failing to add two large numbers, and Distributional has recommended a tool change, prompt change, or guardrail you can implement to address the issue with an assessment of whether the engineering level of effort to accomplish the task.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/demo-fixing-agent-issues-in-production-with-tab-complete-analytics/698d0b75d785a67eddd849d7_5eb4eb19.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;The goal of this approach is to get your team in a virtuous feedback cycle. Users prompt your agent, you collect these traces, Distributional gives you daily insights on what to fix or improve, you implement the change, and you see the performance gains in Distributional over time.&lt;/p&gt;
&lt;h2 id=&quot;get-started&quot;&gt;&lt;a href=&quot;#get-started&quot;&gt;Get started&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The easiest way to &lt;a href=&quot;https://docs.dbnl.com/get-started/quickstart&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;get started&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; is to use a free SaaS demo account to review this example and other examples that we’ve pre-loaded in Distributional. Next, you can install our &lt;a href=&quot;https://docs.dbnl.com/platform/deployment/sandbox&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;sandbox&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; locally on your laptop in ten minutes and run through a tutorial that shows you how to use Distributional for a toy example. Once more familiar with our functionality, you can &lt;a href=&quot;https://docs.dbnl.com/platform/deployment&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;install&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; the full service for free using a Terraform Module or Helm Chart. We are happy to help through any step of this process, so reach out at &lt;a href=&quot;mailto:support@distributional.com&quot;&gt;support@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Lesson notes from the &quot;Hidden Signals in Production AI Logs&quot; session</title><link>https://distributional.com/blog/lesson-notes-from-hidden-signals-in-production-ai-logs-session</link><guid isPermaLink="true">https://distributional.com/blog/lesson-notes-from-hidden-signals-in-production-ai-logs-session</guid><description>Q&amp;A notes on production AI log analytics: analytics versus monitoring versus logging, how behavioral analytics works, and deployment. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Tue, 03 Mar 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Distributional Co-Founder &amp;amp; CEO &lt;a href=&quot;https://www.linkedin.com/in/sc932/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Scott Clark&lt;/a&gt; recently led a &lt;a href=&quot;https://maven.com/p/754cdf/the-hidden-signal-in-production-ai-logs&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;lightning lesson&lt;/a&gt; hosted by &lt;a href=&quot;https://jxnl.co/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Jason Liu&lt;/a&gt; as part of &lt;a href=&quot;https://www.youtube.com/watch?v=FKL918FgxAw&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;his series&lt;/a&gt; of talks helping builders successfully develop, deploy, and scale AI.&lt;/p&gt;
&lt;p&gt;The observability hierarchy for AI systems has three distinct layers: logging and tracing to understand what happened in a specific session, monitoring and evals to determine whether the system is up and passing defined checks, and behavioral analytics to understand patterns across populations of agents and users. This talk focused on the analytics layer of AI observability. You can watch the lesson recording here:&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=QBaSnETJr4Q&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;And here are the lessons that Jason shared with the participants after the session.&lt;/p&gt;
&lt;h2 id=&quot;why-is-analytics-different-from-monitoring-and-logging-for-ai-systems&quot;&gt;&lt;a href=&quot;#why-is-analytics-different-from-monitoring-and-logging-for-ai-systems&quot;&gt;Why is analytics different from monitoring and logging for AI systems?&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Traditional monitoring tells you if your system is up and whether your evals are passing, while logging and tracing let you debug specific sessions. Analytics fills the gap between these two extremes by helping you discover, understand, track, and prioritize hidden behavioral signals across many sessions. Analytics is about discovery, not direct diagnosis. It surfaces candidate patterns that require human judgment to assess importance and guide downstream debugging and product decisions.&lt;/p&gt;
&lt;p&gt;The observability hierarchy for AI systems progresses through stages: first you need basic logging and tracing to see what’s happening, then monitoring to know if the system is up and evals are passing, and finally analytics to understand behavioral patterns across your entire user base.&lt;/p&gt;
&lt;p&gt;Think of it like product analytics for traditional web apps, but instead of tracking users through funnels, you’re tracking agents as the atomic unit. You want to understand how many different agents across many sessions perform specific tasks and what patterns or sub-behaviors emerge.&lt;/p&gt;
&lt;p&gt;This approach completes the AI data flywheel. By knowing what to look for in production data, you can create better evals, better reward functions for fine-tuning or reinforcement learning, and identify specific issues to address through prompt engineering or system improvements. The flywheel works as a continuous loop: observe production behavior, detect emergent patterns, convert those patterns into evals or reward signals, improve system behavior, and repeat.&lt;/p&gt;
&lt;h5 id=&quot;key-takeaway-analytics-bridges-the-gap-between-macro-level-monitoring-and-micro-level-debugging-it-helps-you-find-needles-in-the-haystack-by-analyzing-behavioral-patterns-across-many-ai-sessions-rather-than-examining-each-trace-individually&quot;&gt;&lt;a href=&quot;#key-takeaway-analytics-bridges-the-gap-between-macro-level-monitoring-and-micro-level-debugging-it-helps-you-find-needles-in-the-haystack-by-analyzing-behavioral-patterns-across-many-ai-sessions-rather-than-examining-each-trace-individually&quot;&gt;Key takeaway: Analytics bridges the gap between macro-level monitoring and micro-level debugging. It helps you find needles in the haystack by analyzing behavioral patterns across many AI sessions rather than examining each trace individually.&lt;/a&gt;&lt;/h5&gt;
&lt;h2 id=&quot;how-does-distributionals-approach-to-behavioral-analytics-actually-work&quot;&gt;&lt;a href=&quot;#how-does-distributionals-approach-to-behavioral-analytics-actually-work&quot;&gt;How does Distributional’s approach to behavioral analytics actually work?&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The system operates as an unsupervised data flywheel. First, trace data from your agentic system gets enriched with behavioral signals. These can be LLM-as-judge evals, classic NLP statistical measures, or signals from other tools. The goal is adding as much enriched signal as possible to whatever data your app naturally produces. Unsupervised methods matter because you cannot label or define failures you do not yet know exist.&lt;/p&gt;
&lt;p&gt;Every trace gets a behavioral vector in high-dimensional space representing what actually happened. The analysis phase looks at distributions of these vectors across many traces to pull out sub-pockets that represent infrequent behaviors or patterns correlated with cost, latency, or quality issues.&lt;/p&gt;
&lt;p&gt;These subclusters get fed into an LLM backend to generate insights. The system uses cheaper, faster models for high-volume behavioral evaluation, then more capable mid-weight models with reasoning capabilities for final insight generation and fix recommendations.&lt;/p&gt;
&lt;p&gt;The philosophy is “many weak signals are better than a single strong signal.” Instead of trying to create one perfect eval for quality, you combine multiple signals around frustration, tone, verbosity, reading level, and other dimensions. Through clustering, you extract the strong signal for the performance you actually care about.&lt;/p&gt;
&lt;p&gt;Distributional provides a guided analytics experience designed to collapse weeks of manual data analysis into fast, structured triage. Instead of spending a week doing data science yourself, insights are presented to you so you can quickly assess whether an issue matters, investigate it with specific evidence, and track whether your fixes worked.&lt;/p&gt;
&lt;h5 id=&quot;key-takeaway-distributional-combines-unsupervised-learning-on-behavioral-vectors-with-llm-analysis-to-automatically-discover-issues-you-didnt-know-to-look-for-complete-with-evidence-and-suggested-fixes&quot;&gt;&lt;a href=&quot;#key-takeaway-distributional-combines-unsupervised-learning-on-behavioral-vectors-with-llm-analysis-to-automatically-discover-issues-you-didnt-know-to-look-for-complete-with-evidence-and-suggested-fixes&quot;&gt;Key takeaway: Distributional combines unsupervised learning on behavioral vectors with LLM analysis to automatically discover issues you didn’t know to look for, complete with evidence and suggested fixes.&lt;/a&gt;&lt;/h5&gt;
&lt;h2 id=&quot;what-types-of-issues-does-behavioral-analytics-typically-discover&quot;&gt;&lt;a href=&quot;#what-types-of-issues-does-behavioral-analytics-typically-discover&quot;&gt;What types of issues does behavioral analytics typically discover?&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;One common pattern is redundant and inefficient tool usage. The demo showed an agent making duplicate Google Maps searches for the same information within a single session. The system detected this pattern, provided specific evidence with exact traces, and suggested fixes ranging from simple prompt changes to implementing caching systems with guardrails.&lt;/p&gt;
&lt;p&gt;Expansion-related failures happen frequently when systems scale. A company might develop their agent with a small group where it scores well on evals, then roll it out to the rest of the company or internationally and everything breaks. Users in different regions ask questions differently. What one country calls “parental leave” another calls “maternity leave,” triggering guardrails inappropriately.&lt;/p&gt;
&lt;p&gt;Combinatorial complexity issues emerge as you add more tools. An agent might work well with four MCP servers, but when you add both GitHub and Linear, the system fails in non-linear ways because it can’t distinguish between issues and tickets. The order in which MCP tools load can cause odd behavior.&lt;/p&gt;
&lt;p&gt;Edge cases in production that offline evals miss are surprisingly common. Companies report rare errors like 403 responses in 0.2% of traffic that weren’t caught in testing. These rough edges accumulate across agentic systems with their directed graphs of tool calls.&lt;/p&gt;
&lt;p&gt;Another pattern is agents getting stuck in loops when they can’t return something, calling the same tool repeatedly. Since these directed graphs aren’t always acyclic, they can go off the rails quickly in ways that are hard to anticipate.&lt;/p&gt;
&lt;h5 id=&quot;key-takeaway-most-issues-are-behavioral-patterns-that-emerge-only-in-production-at-scale-these-include-inefficient-tool-usage-unexpected-user-interaction-patterns-edge-cases-in-tool-calling-sequences-and-failures-that-only-appear-with-certain-user-segments-or-geographical-regions&quot;&gt;&lt;a href=&quot;#key-takeaway-most-issues-are-behavioral-patterns-that-emerge-only-in-production-at-scale-these-include-inefficient-tool-usage-unexpected-user-interaction-patterns-edge-cases-in-tool-calling-sequences-and-failures-that-only-appear-with-certain-user-segments-or-geographical-regions&quot;&gt;Key takeaway: Most issues are behavioral patterns that emerge only in production at scale. These include inefficient tool usage, unexpected user interaction patterns, edge cases in tool-calling sequences, and failures that only appear with certain user segments or geographical regions.&lt;/a&gt;&lt;/h5&gt;
&lt;h2 id=&quot;how-should-companies-think-about-the-relationship-between-evals-and-production-analytics&quot;&gt;&lt;a href=&quot;#how-should-companies-think-about-the-relationship-between-evals-and-production-analytics&quot;&gt;How should companies think about the relationship between evals and production analytics?&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Offline evals alone are insufficient because it’s impossible to anticipate everything that can happen in production. Users will use your system in different ways than you expect, and the foundational models underneath are non-stationary and changing continuously.&lt;/p&gt;
&lt;p&gt;Getting 100 percent on your evals doesn’t mean you’ll get 100 percent in the real world. You can pass a kindergarten math test with perfect scores, but that tells you nothing about real-world performance. Just like you could overfit a random forest to get 100 percent on a test set years ago, you can create evals that don’t capture actual system behavior.&lt;/p&gt;
&lt;p&gt;The value of production analytics is discovering unknown unknowns that you can then convert into known issues. When you find a pattern in production, you can create an eval to catch that behavior going forward or use it as a reward function in reinforcement learning.&lt;/p&gt;
&lt;p&gt;This isn’t a replacement for evals. It’s part of the flywheel. You observe what happens in the real world, discover new failure modes, add those to your eval suite, and continuously improve. The only way to see emergent behaviors is by observing production.&lt;/p&gt;
&lt;p&gt;For agentic systems, binary pass or fail percentages don’t even make sense. You’re operating in a continuous behavioral space that you’re trying to guide your agent through. Many weak behavioral signals combined give you better understanding than any single strong eval could provide.&lt;/p&gt;
&lt;h5 id=&quot;key-takeaway-production-analytics-and-offline-evals-work-together-in-a-flywheel-analytics-discovers-the-unknown-issues-that-become-tomorrows-evals-creating-continuous-improvement-rather-than-one-time-validation&quot;&gt;&lt;a href=&quot;#key-takeaway-production-analytics-and-offline-evals-work-together-in-a-flywheel-analytics-discovers-the-unknown-issues-that-become-tomorrows-evals-creating-continuous-improvement-rather-than-one-time-validation&quot;&gt;Key takeaway: Production analytics and offline evals work together in a flywheel. Analytics discovers the unknown issues that become tomorrow’s evals, creating continuous improvement rather than one-time validation.&lt;/a&gt;&lt;/h5&gt;
&lt;h2 id=&quot;when-do-companies-actually-need-behavioral-analytics-for-their-ai-systems&quot;&gt;&lt;a href=&quot;#when-do-companies-actually-need-behavioral-analytics-for-their-ai-systems&quot;&gt;When do companies actually need behavioral analytics for their AI systems?&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The need escalates as system complexity increases. If you’re building a simple chatbot with one-shot questions, basic monitoring might suffice. But as you move from chatbots to RAG systems to actual agents performing work, you need to climb up the observability hierarchy.&lt;/p&gt;
&lt;p&gt;Companies building internal AskHR or AskIT-style systems increasingly add tooling so the system can fix problems directly rather than just pointing users to documentation. When you ask about PTO, it processes the request instead of linking to a website. These systems become combinatorially complex extremely rapidly. Even the router deciding which tool to call becomes interesting to analyze.&lt;/p&gt;
&lt;p&gt;Regulated industries particularly value the on-premise, secure-first approach where you own the data and models. This matters for companies that can’t send production data to external services.&lt;/p&gt;
&lt;p&gt;The analytics become essential when you care about performance, quality, and behavioral understanding at scale. If you’re just trying to get something to work initially, you don’t need this yet. But when you want to scale and be best in class, you need to understand behavioral patterns.&lt;/p&gt;
&lt;p&gt;One clear indicator is when you start seeing unexplainable degradation. A team focused on making email search excellent discovered through observability that 30 percent of search queries were actually for photos, screenshots of purchase orders taken on phones. They spent a month optimizing the wrong thing because they lacked visibility into actual usage patterns.&lt;/p&gt;
&lt;h5 id=&quot;key-takeaway-you-need-behavioral-analytics-when-moving-from-prototype-to-production-scale-when-system-complexity-increases-beyond-simple-chatbots-when-youre-in-regulated-industries-requiring-on-premise-solutions-or-when-you-need-to-understand-actual-user-behavior-rather-than-assumed-behavior&quot;&gt;&lt;a href=&quot;#key-takeaway-you-need-behavioral-analytics-when-moving-from-prototype-to-production-scale-when-system-complexity-increases-beyond-simple-chatbots-when-youre-in-regulated-industries-requiring-on-premise-solutions-or-when-you-need-to-understand-actual-user-behavior-rather-than-assumed-behavior&quot;&gt;Key takeaway: You need behavioral analytics when moving from prototype to production scale, when system complexity increases beyond simple chatbots, when you’re in regulated industries requiring on-premise solutions, or when you need to understand actual user behavior rather than assumed behavior.&lt;/a&gt;&lt;/h5&gt;
&lt;h2 id=&quot;how-does-data-get-into-distributional-and-whats-the-deployment-model&quot;&gt;&lt;a href=&quot;#how-does-data-get-into-distributional-and-whats-the-deployment-model&quot;&gt;How does data get into Distributional and what’s the deployment model?&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The system integrates through multiple paths. Many agent frameworks already have OpenTelemetry built in, so you can route traces to Distributional the same way you’d route to Datadog or CloudWatch using a write-once, send-many approach.&lt;/p&gt;
&lt;p&gt;Companies with robust ETL pipelines can ingest data through Parquet files, sit on top of Iceberg tables, or use SQL ingestion if events exist within larger databases. The fundamental goal is taking the richest possible version of your data. If you only have inputs and outputs, some insight is possible, but adding tracing data, user feedback, session-level events, and evals from other tools enables deeper analysis.&lt;/p&gt;
&lt;p&gt;The product deploys as open source and free. It’s distributed as a Kubernetes cluster for full deployment or a K3D cluster within a single Docker image for the sandbox version. The sandbox can run on your laptop and be operational in under an hour.&lt;/p&gt;
&lt;p&gt;Distributional never sees your underlying data. Everything runs on-premise in your cloud or bare metal infrastructure. This architecture matters for companies in regulated industries or those with strict data governance requirements.&lt;/p&gt;
&lt;p&gt;The system is agnostic to your agent backend, cloud provider, and existing tooling. It bolts on top of whatever traces you’re already writing for logging and monitoring.&lt;/p&gt;
&lt;h5 id=&quot;key-takeaway-integration-happens-through-standard-observability-protocols-like-opentelemetry-with-flexible-ingestion-options-for-different-data-formats-the-on-premise-deployment-model-means-your-data-stays-in-your-infrastructure-while-you-gain-analytics-capabilities&quot;&gt;&lt;a href=&quot;#key-takeaway-integration-happens-through-standard-observability-protocols-like-opentelemetry-with-flexible-ingestion-options-for-different-data-formats-the-on-premise-deployment-model-means-your-data-stays-in-your-infrastructure-while-you-gain-analytics-capabilities&quot;&gt;Key takeaway: Integration happens through standard observability protocols like OpenTelemetry, with flexible ingestion options for different data formats. The on-premise deployment model means your data stays in your infrastructure while you gain analytics capabilities.&lt;/a&gt;&lt;/h5&gt;
&lt;h2 id=&quot;whats-the-philosophy-behind-making-analytics-a-priority-in-ai-development&quot;&gt;&lt;a href=&quot;#whats-the-philosophy-behind-making-analytics-a-priority-in-ai-development&quot;&gt;What’s the philosophy behind making analytics a priority in AI development?&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;History shows that every major paradigm shift in software, web, microservices, mobile, follows the same pattern. Building comes first, then logging to see what happens, then monitoring, and finally analytics to squeeze out maximum value and create amazing products.&lt;/p&gt;
&lt;p&gt;The mindset shift is from “I don’t see any problems, therefore they don’t exist” to “I want to know about unknown unknowns.” Stopping testing for a disease doesn’t mean the disease is cured. You need active investigation to understand what’s actually happening.&lt;/p&gt;
&lt;p&gt;The ideal customer is a product owner who treats their product as something they care about and want to improve, not a checklist where as long as PagerDuty isn’t alerting, everything’s fine. As agentic systems provide more enterprise value, more people will need to take on that agency and care.&lt;/p&gt;
&lt;p&gt;The goal is helping you play whack-a-mole with issues. There’s always another problem that comes up, whether in parenting, fighting disease, or managing AI systems. Analytics provides the flashlight so you’re not wandering in the dark or only looking at your evals while ignoring the rest of the room.&lt;/p&gt;
&lt;p&gt;You optimize what you measure, but you can only measure what you know to look for. Analytics helps you see more, measure more, and hopefully optimize for the right things rather than local maxima in your eval suite.&lt;/p&gt;
&lt;p&gt;The final message: don’t be pigeonholed into only the specific evals you’re looking at. Look around and try to find as many unknowns as possible, because the big solutions are usually behavioral investments in tooling and understanding rather than tweaking words in system prompts.&lt;/p&gt;
&lt;h5 id=&quot;key-takeaway-analytics-is-inevitable-in-the-ai-development-lifecycle-the-question-isnt-whether-youll-need-it-but-when-youll-start-investing-in-understanding-behavioral-patterns-to-complement-your-monitoring-and-debugging-tools&quot;&gt;&lt;a href=&quot;#key-takeaway-analytics-is-inevitable-in-the-ai-development-lifecycle-the-question-isnt-whether-youll-need-it-but-when-youll-start-investing-in-understanding-behavioral-patterns-to-complement-your-monitoring-and-debugging-tools&quot;&gt;Key takeaway: Analytics is inevitable in the AI development lifecycle. The question isn’t whether you’ll need it, but when you’ll start investing in understanding behavioral patterns to complement your monitoring and debugging tools.&lt;/a&gt;&lt;/h5&gt;
&lt;h2 id=&quot;get-started&quot;&gt;&lt;a href=&quot;#get-started&quot;&gt;Get started&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional is a free, open, and installable platform for agent analytics. &lt;a href=&quot;https://docs.dbnl.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Try it today&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; and quickly learn how it complements your existing agent observability stack. We are also always happy to learn more about your use case and enterprise needs, so reach out to &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;contact@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Analytics-driven agent A/B testing</title><link>https://distributional.com/blog/analytics-driven-agent-a-b-testing</link><guid isPermaLink="true">https://distributional.com/blog/analytics-driven-agent-a-b-testing</guid><description>End-to-end agent A/B testing: discover an issue from behavioral signals, root-cause it, fix it, and verify with distribution comparison. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 18 Feb 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;“First, it was hard to find the issue with our AI agent. We saw feedback dip and then usage drop, but didn’t know why,” explained an AI engineer at a large insurance company. “Next, it took our engineering team an entire week (1 engineering month) of work to actually diagnose the root cause of the issue and make the change.” They continued, “Finally, we weren’t actually sure that the change fixed the issue, and still aren’t today – we saw feedback improve and usage has slowly started to crawl back up, but both may be disconnected from the actual problem.”&lt;/p&gt;
&lt;p&gt;We’ve heard the same story from dozens of AI product teams: it is hard to discover and diagnose issues when your agent is in production, and can be even harder to confirm the fix is working. These stories motivated our approach to adaptive analytics. In this &lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/ab_test_example&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;example&lt;/a&gt;, we show how you can solve each of these problems with Distributional’s product (&lt;a href=&quot;https://docs.dbnl.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;DBNL&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;) to have a much more rapid cycle of fixes and improvements for your production AI agents. To jump straight to the example and start using it, navigate here: &lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/ab_test_example&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://github.com/dbnlAI/examples/tree/main/ab_test_example&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;In this short demo, you will learn how to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Automatically discover new issues with your AI agent&lt;/li&gt;
&lt;li&gt;Rapidly analyze the issue to uncover the root cause&lt;/li&gt;
&lt;li&gt;A/B test a fix to be confident it solved the issue&lt;/li&gt;
&lt;li&gt;Track the fix to make sure it doesn’t reappear&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;scenario&quot;&gt;&lt;a href=&quot;#scenario&quot;&gt;Scenario&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;To simplify this example, we built an LLM-based calculator agent that calls tools to add, subtract, multiply, and divide numbers. We generated a wide variety of math strings and ran these as queries to this calculator. We then intentionally introduced a bug in one of the tools for DBNL to automatically discover. We logged these traces and augmented them with a variety of properties:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Expected value: Value we would expect when performing the calculation&lt;/li&gt;
&lt;li&gt;Simulated Feedback: Thumbs up or down from the user and feedback text&lt;/li&gt;
&lt;li&gt;Version: Extract the version of the agent to facilitate A/B testing of fixes&lt;/li&gt;
&lt;li&gt;Absolute error: Difference between expected and actual output&lt;/li&gt;
&lt;li&gt;Cost: Tokens and estimated cost&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We ran this calculator for 7 days with synthetically generated queries, at which point DBNL began producing daily insights on cost, quality, speed, and usage of the agent. (Typically AI product teams backfill with historical data to get started with DBNL much faster.)&lt;/p&gt;
&lt;h2 id=&quot;understand-usage&quot;&gt;&lt;a href=&quot;#understand-usage&quot;&gt;Understand usage&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;DBNL makes sense of production AI logs with &lt;a href=&quot;https://docs.dbnl.com/workflow/adaptive-analytics-workflow&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;analytics&lt;/a&gt; that tell a deeper, more complete story on cost, speed, quality, and usage of the AI agent.&lt;/p&gt;
&lt;p&gt;At the top of the DBNL &lt;a href=&quot;https://docs.dbnl.com/workflow/dashboards&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Dashboard&lt;/a&gt; is a collection of high-level KPIs. In this case, we track user feedback as our product KPI and an indicator of whether someone is getting value out of the calculator or not. Immediately, we see something that may be worth investigating, which is a dip in average user feedback score and a lower score (3.76 on a Likert scale of 1 - 5) than we expected.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941e83e17522b79e5ae8ebf_image9.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;DBNL automatically computes standard and LLM-as-judge &lt;a href=&quot;https://docs.dbnl.com/workflow/metrics&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Metrics&lt;/a&gt; like Output Relevancy and User Frustration to give a sense of the quality of user experience over time.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941e85ebc314c8bade4d9db_image6.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;DBNL also tracks tool calls for an overarching sense of how the agent system is working so you can get a quick sense of whether this is in line with expectations.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941e88277763134bb17c89f_image16.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;And DBNL visualizes the tool call sequence flow to provide deeper understanding of how the agent is using these tools and more intuitive debugging of issues around agent performance. This call sequence flow can also be viewed, filtered, and grouped by a variety of attributes so you can get a quick sense of whether the agent is performing as expected.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6931e84c27ea89d09016bc27_image2.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;discover-an-issue-from-a-signal&quot;&gt;&lt;a href=&quot;#discover-an-issue-from-a-signal&quot;&gt;Discover an issue from a signal&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;We already passively saw a signal that suggested a potential issue to investigate. Feedback scores seemed to be lower than expected, and lower than recent averages. In this same &lt;a href=&quot;https://docs.dbnl.com/workflow/dashboards&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Dashboard&lt;/a&gt;, you can go deeper on feedback scores by looking at time series and distributions. For quick investigation, you can also directly click into logs or explorer, which we will show in the next section.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941e8bbffda954cba5a3eda_image15.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;DBNL also offers a more proactive way to discover issues that also serves to more rapidly jumpstart investigation. Every day, DBNL produces &lt;a href=&quot;https://docs.dbnl.com/workflow/insights&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Insights&lt;/a&gt; that are human readable summaries of signals the product found hidden in your production AI logs. These &lt;a href=&quot;https://docs.dbnl.com/workflow/insights&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Insights&lt;/a&gt; can be found in your Dashboard or through &lt;a href=&quot;https://docs.dbnl.com/configuration/notification-connections&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Notification Connections&lt;/a&gt; that you configure. In this case, we immediately see the Incorrect or Malformed Tool Output issue is high severity and worth investigating, listed first in the Insights from this day.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941e8e548dd6220bc299fcd_image19.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;perform-rapid-root-cause-analysis&quot;&gt;&lt;a href=&quot;#perform-rapid-root-cause-analysis&quot;&gt;Perform rapid root cause analysis&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Alongside these Insights that discover issues, DBNL provides an explanation and evidence to jumpstart root cause analysis. We click into the Incorrect and Malformed Tool Output and see a simple description of the issue and a selection of relevant log examples. Immediately, we see that the issue appears to be with addition. When asked to add two numbers together, this agent is responding incorrectly.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6931e863a948887ed921cfd7_image4.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;In this case, we likely wouldn’t need additional evidence, as the root cause analysis is clear. For more complex cases, however, it may help to directly investigate the logs. DBNL provides a click through to a filtered set of the most relevant &lt;a href=&quot;https://docs.dbnl.com/workflow/logs&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Logs&lt;/a&gt; responsible for the issue.&lt;/p&gt;
&lt;p&gt;You can see in these logs that the absolute error is abnormally high.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941e98d6abe1b5341685e6b_image4.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;You also see that feedback is particularly negative and to the point on the low quality of the math output, confirming the issue.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941e9a887c92576ed6ac80c_image11.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;In some cases, it may also help to visualize the issue. DBNL has an &lt;a href=&quot;https://docs.dbnl.com/workflow/explorer&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Explorer&lt;/a&gt; that provides filterable charts and graphs to bootstrap this type of quick analysis.&lt;/p&gt;
&lt;p&gt;When we filter by this expected issue, we get a &lt;a href=&quot;https://docs.dbnl.com/workflow/explorer#segment-comparison&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Segment Comparison&lt;/a&gt; that bifurcates the data into A (where the addition tool is called, pink) and B (the entire dataset, orange) segments. Explorer shows that Segment A has much higher absolute error count than Segment B, we are more likely to see an error when addition is called than from the general population of all tool call chains.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941e9d0f9e90016ad7e29e9_image3.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Progressing further down the &lt;a href=&quot;https://docs.dbnl.com/workflow/explorer&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Explorer&lt;/a&gt; page, you see Segment A has a much lower average feedback score than Segment B.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941e9eda8ba9823801112fe_image14.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;fix-the-issue&quot;&gt;&lt;a href=&quot;#fix-the-issue&quot;&gt;Fix the issue&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;DBNL provides suggestions to fix an issue. Fixes happen off of the DBNL platform alongside the root cause analysis that assists in guiding the fix.&lt;/p&gt;
&lt;p&gt;Directly attached to the insight, DBNL proposes a set of potential fixes. In this case, DBNL recommends a tool change, guardrail, or prompt adjustment, with different tradeoffs and an estimated effort to apply the change.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941ea0e7b32f5c2b3dcde8f_image5.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;We take DBNL’s advice, and start by checking the tool. Immediately, we see an error where our addition tool multiplies two numbers instead of adding them.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;def add_two_numbers(a: float, b: float) -&amp;gt; dict:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &amp;quot;&amp;quot;&amp;quot;Returns the sum of two numbers by adding them together&amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    # We intentionally introduce a bug where it gives the wrong answer because we accidentally typed `*` instead of `+`&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    return {&amp;quot;status&amp;quot;: &amp;quot;ok&amp;quot;, &amp;quot;result&amp;quot;: a * b}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Alongside DBNL’s recommendations for fixes, DBNL also suggests segments you can track so that you can keep tabs on an issue you want to triage later, ensure a fix actually addresses an issue, and ensure that resolved issues don’t reemerge. By adding a new &lt;a href=&quot;https://docs.dbnl.com/workflow/segments&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Segment&lt;/a&gt;, DBNL will add this to the &lt;a href=&quot;https://docs.dbnl.com/workflow/dashboards#segments-dashboard&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Segments Dashboard&lt;/a&gt; and provide &lt;a href=&quot;https://docs.dbnl.com/workflow/insights&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Insights&lt;/a&gt; on it going forward.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941ea5324466251b9792598_image7.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;ab-test-the-fix&quot;&gt;&lt;a href=&quot;#ab-test-the-fix&quot;&gt;A/B test the fix&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;When deploying a fix, AI product teams often want to A/B test the solution. DBNL offers a convenient way to analyze this A/B test to confirm the fix. In this example we will roll out a fix as a 50/50 A/B test.&lt;/p&gt;
&lt;p&gt;For this calculator, we are able to extract the cohort each trace belongs to by tracking the product version as an attribute. When we made the fix, we created two versions of the agent, v0 to v1, and we can now filter on this attribute in &lt;a href=&quot;https://docs.dbnl.com/workflow/explorer&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Explorer&lt;/a&gt; and compare Segment A (pink, v0) to Segment B (orange, v1).&lt;/p&gt;
&lt;p&gt;We’d expect to see a reduction in absolute error for v1 (orange), and we see this confirmed – error is now zero for v1.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941ea74f3e2034dd0b9ea15_image13.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Conversely, we’d expect to see average feedback score jump for v1 (orange) compared to v0 (pink). For v1 (orange), we see that feedback is universally positive, with an average score of 5.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941ea91eca834072a09eb64_image18.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;We even see this fix showing up in Insights, where a positive distribution drift was detected.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941eaaff3e2034dd0b9f1df_image12.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;track-the-fix&quot;&gt;&lt;a href=&quot;#track-the-fix&quot;&gt;Track the fix&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Regardless of whether you A/B test the fix or not, you will want to track it to confirm the change is working as expected and the issue has been resolved. You can do this in DBNL.&lt;/p&gt;
&lt;p&gt;First, we see that DBNL passively tracks this and reflects the change in the KPI at the top of the &lt;a href=&quot;https://docs.dbnl.com/workflow/dashboards&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Dashboard&lt;/a&gt;. Average feedback score has jumped back up.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941eacd23e187459b1b3298_image10.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;You can also see this reflected in the feedback score time series, with all recent days showing 5 out of 5 for user feedback.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941eae45d715f2666ef2922_image8.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Jumping back to Explorer, we can track these changes in metrics as well. We see absolute error dropping to zero in this chart as one example.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/analytics-driven-agent-a-b-testing/6941eaff31c70c6fd8cac161_image2.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;improve-ai-with-signals-from-production-logs&quot;&gt;&lt;a href=&quot;#improve-ai-with-signals-from-production-logs&quot;&gt;Improve AI with signals from production logs&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;In this example, we show you how to use DBNL to discover, investigate, a/b test, and track a fix to an AI calculator agent in production. This calculator agent is a simplified example, but all of the same features in DBNL can be used for even the most complex AI agents receiving millions of requests per day.&lt;/p&gt;
&lt;p&gt;The key to this workflow starts with finding signals for what to fix or improve from the AI production logs themselves. When DBNL—rather than bespoke data analysis—does this automatically, it ensures you don’t miss a signal you didn’t know to look or eval for, which could end up being a big issue. If we didn’t catch the basic addition error, nobody would end up using this calculator over time. Similarly, grounding this analysis with direct evidence from logs makes it much easier to run A/B tests and confirm the change after it is made. And instantiating this full workflow in DBNL gives the entire team a central source of information on what is and isn’t working with their complex AI products.&lt;/p&gt;
&lt;p&gt;All of this is useful for this calculator example, but becomes even more critical as the complexity of the AI agent increases or scale of usage grows. Run this &lt;a href=&quot;https://github.com/dbnlAI/examples&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;example&lt;/a&gt; to get familiar with DBNL or &lt;a href=&quot;https://docs.dbnl.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;install&lt;/a&gt; the full DBNL service for free in your environment to get started on your own agent.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>DBNL fits your agent data</title><link>https://distributional.com/blog/dbnl-fits-your-agent-data</link><guid isPermaLink="true">https://distributional.com/blog/dbnl-fits-your-agent-data</guid><description>Agent data ingestion: OTel traces versus SDK logs versus SQL pulls, the DBNL Semantic Convention, and concrete adapters. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 11 Feb 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Distributional’s product (DBNL) is designed to fit into your existing production AI stack. This means you can use DBNL without fear of lock-in and with minimal engineering overhead. Our &lt;a href=&quot;https://docs.dbnl.com/configuration/model-connections&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Model Connections&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; and &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Data Connections&lt;/a&gt; make it easy to integrate DBNL with whatever LMs you already have in place for your AI products and leverage the data you already have, wherever it may be, to start performing behavioral analytics on your agents right away.&lt;/p&gt;
&lt;p&gt;This post focuses on &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Data Connections&lt;/a&gt; and, more specifically, examples for fitting your data to the &lt;a href=&quot;https://docs.dbnl.com/configuration/dbnl-semantic-convention&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;DBNL Semantic Convention&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;data-connections&quot;&gt;&lt;a href=&quot;#data-connections&quot;&gt;Data Connections&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional’s &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Data Connections&lt;/a&gt; include support for three approaches for DBNL to ingest production AI logs. &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections/otel-trace-ingestion&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Open Telemetry (OTEL) Trace Ingestion&lt;/a&gt; publishes Open Telemetry traces directly to DBNL as your agent runs. &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections/sdk-log-ingestion&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;SDK Log Ingestion&lt;/a&gt; pushes data manually or as part of a daily orchestration job using the Python SDK. &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections/sdk-log-ingestion&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;SQL Integration Ingestion&lt;/a&gt; pulls data from a SQL table into DBNL on a schedule.&lt;/p&gt;
&lt;p&gt;We’ve also written a few adapters to conform your traces and spans to our &lt;a href=&quot;https://docs.dbnl.com/configuration/dbnl-semantic-convention&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Semantic Convention&lt;/a&gt; that will work with our SDK or your OTEL connections, which are included in this post below.&lt;/p&gt;
&lt;h2 id=&quot;semantic-convention&quot;&gt;&lt;a href=&quot;#semantic-convention&quot;&gt;Semantic Convention&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;DBNL ingests data using traces produced by telemetry frameworks with different semantic conventions as well as tabular Logs with a user defined format. To compute &lt;a href=&quot;https://docs.dbnl.com/workflow/metrics&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Metrics&lt;/a&gt; and derive &lt;a href=&quot;https://docs.dbnl.com/workflow/metrics&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Insights&lt;/a&gt; consistently across different data ingestion formats, we define a &lt;a href=&quot;https://docs.dbnl.com/configuration/dbnl-semantic-convention&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Semantic Convention&lt;/a&gt; for the data uploaded to DBNL to get the richest possible analytics from your data.&lt;/p&gt;
&lt;p&gt;If you are using &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections/otel-trace-ingestion&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;OTEL Trace Ingestion&lt;/a&gt; and the &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections/otel-trace-ingestion&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;OpenInference&lt;/a&gt; semantic convention, this happens automatically. If you are using &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections/sdk-log-ingestion&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;SDK Log Ingestion&lt;/a&gt; or &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections/sql-integration-ingestion&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;SQL Integration Ingestion&lt;/a&gt;, you need to ensure that your Column names adhere to our semantic convention for best results. In this post, we’ll explain a few adapters to conform your standard to fit with DBNL’s &lt;a href=&quot;https://docs.dbnl.com/configuration/dbnl-semantic-convention&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Semantic Convention&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;adapters-for-dbnls-semantic-convention&quot;&gt;&lt;a href=&quot;#adapters-for-dbnls-semantic-convention&quot;&gt;Adapters for DBNL’s Semantic Convention&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Below is a chart that shows pathways from your AI agent to DBNL &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Data Connections&lt;/a&gt; that maps to our &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Semantic Convention&lt;/a&gt;. We’ve also highlighted the recommended path.&lt;/p&gt;
&lt;p&gt;You’ll notice that each of these pathways uses the OpenTelemetry standard. One of the pathways pushes traces directly from OTEL into DBNL, but we actually recommend instead taking steps to pass the data through our SDK so you can augment the data with rich fields potentially unavailable at span creation like user feedback, session information, or other data you can join after the fact. This richer data enables DBNL to produce more compelling analytics and provide deeper insight on AI agent usage, cost, quality, and speed.&lt;/p&gt;
&lt;p&gt;Here is a quick overview of each pathway with a link to an example based on what we recommend given your circumstance.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/dbnl-fits-your-agent-data/692f3e082837d204b55e0eec_Screenshot-2025-12-02-at-11.28.52-AM.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;And here is a short description of each:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/adk_calculator_sdk_from_otel&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Agent =&amp;gt; OTEL =&amp;gt; Traces JSONL =&amp;gt; Augment =&amp;gt; SDK =&amp;gt; DBNL&lt;/a&gt;: Use a local OTEL collector to write raw spans to file in the OpenInference semantic convention using resourceSpans, which are then easily augmented and uploaded to DBNL via the SDK. These can be converted to logs using &lt;code&gt;flatten_otlp_traces_data&lt;/code&gt; and reported using &lt;code&gt;dbnl.log&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/adk_calculator_otel_direct&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Agent =&amp;gt; OTEL =&amp;gt; DBNL&lt;/a&gt;: Send OTEL traces directly to DBNL without augmentation. This is the simplest path to getting data into DBNL and requires the fewest intermediate steps, storage, and computation. It requires the agent to send all required and useful information via OTEL directly, which may require extra instrumentation.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/adk_calculator_sdk_from_json&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Agent =&amp;gt; OTEL =&amp;gt; Raw JSONL =&amp;gt; SDK =&amp;gt; DBNL&lt;/a&gt;: Extract all of the information needed for the DBNL semantic convention, put it into a local jsonl file, load into a dataframe, and upload it to DBNL via the SDK.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/dbnlAI/examples/tree/main/sdk_from_langfuse_export&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Agent =&amp;gt; OTEL =&amp;gt; Langfuse =&amp;gt; Export File =&amp;gt; Augment =&amp;gt; SDK =&amp;gt; DBNL&lt;/a&gt;: Convert Langfuse trace and observation export files into the DBNL &lt;a href=&quot;https://docs.dbnl.com/configuration/dbnl-semantic-convention&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Semantic Convention&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Finally, here is a graphic that may be easier to follow:&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/dbnl-fits-your-agent-data/692f3e5dd971903390c8aba2_image1.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;get-started-today&quot;&gt;&lt;a href=&quot;#get-started-today&quot;&gt;Get started today&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This platform is designed to fit with your existing stack. It is deployed in your environment and no data leaves, so it is fully secure. It comes with a full set of enterprise features for administration, authentication, and data security. And it is built to scale with minimal engineering overhead. These adapters are an example of how we are constantly evolving our product to be even easier to use with your existing AI agent setup.&lt;/p&gt;
&lt;p&gt;Distributional’s full service is open and free to use. &lt;a href=&quot;https://docs.dbnl.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Install DBNL today&lt;/a&gt; to experience these enterprise features yourself.&lt;/p&gt;
&lt;p&gt;We are also always happy to learn more about your use case and enterprise needs, so reach out to &lt;a href=&quot;mailto:nick-dbnl@distributional.com&quot;&gt;nick-dbnl@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with any questions.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Integrating DBNL with NVIDIA NeMo Agent Toolkit</title><link>https://distributional.com/blog/integrating-dbnl-with-the-nvidia-nemo-agent-toolkit</link><guid isPermaLink="true">https://distributional.com/blog/integrating-dbnl-with-the-nvidia-nemo-agent-toolkit</guid><description>Seven-step tutorial for integrating DBNL with the NVIDIA NeMo Agent Toolkit, from install to analyzing traces. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 21 Jan 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Distributional’s adaptive analytics platform empowers AI product teams to have confidence in AI &lt;em&gt;behavior&lt;/em&gt;—the interplay and correlations between users, context, tools, models, and metrics. Distributional’s &lt;a href=&quot;https://www.distributional.com/product&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;product&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; (DBNL) is composed of three components: platform, data pipeline, and analytics workflow. We designed our platform to fit seamlessly with your existing AI stack so there is minimal overhead to implementing our product, and no lock-in when using it.&lt;/p&gt;
&lt;p&gt;Part of this approach to our platform includes investing in integrations to make our product fit even more cleanly with partner products. This approach is especially important in the context of our &lt;a href=&quot;https://docs.dbnl.com/configuration/data-connections&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Data Connections&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; that make it easier for Distributional to ingest semantically relevant data so our data pipeline can produce compelling insights that guide the analytics workflow.&lt;/p&gt;
&lt;h2 id=&quot;nvidia-nemo-agent-toolkit-and-dbnl-integration&quot;&gt;&lt;a href=&quot;#nvidia-nemo-agent-toolkit-and-dbnl-integration&quot;&gt;NVIDIA NeMo Agent Toolkit and DBNL integration&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This issue crops up frequently in the context of agents. Agents typically have complex workflows with many tool calls, tasks, data sources, and decision points. This complexity can be hard to parse, but the richness of this data presents an opportunity for more useful and insightful analytics—especially as use of these agents scales.&lt;/p&gt;
&lt;p&gt;NVIDIA &lt;a href=&quot;https://github.com/NVIDIA/NeMo-Agent-Toolkit/tree/develop&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NeMo Agent Toolkit&lt;/a&gt;  is a collection of tools that make it easy to build an agent regardless of which agent framework you use (or even if you roll your own). It includes integrations with popular tools like MCP and Google’s ADK, as well as Function Groups and Automatic Hyperparameter Tuning. It also powers observability to provide visibility into traces as usage scales in production. Consistent with the full complement of NVIDIA &lt;a href=&quot;https://github.com/NVIDIA/NeMo-Agent-Toolkit/tree/develop&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;NeMo&lt;/a&gt; tools, the NeMo Agent Toolkit is designed to help AI product teams efficiently design, optimize, and scale their agents on GPUs.&lt;/p&gt;
&lt;p&gt;This is why Distributional built an &lt;a href=&quot;https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/develop/docs/source/run-workflows/observe/observe-workflow-with-dbnl.md&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;integration&lt;/a&gt; with NeMo Agent Toolkit to make it easy for any AI product team to ingest NeMo Agent Toolkit traces in DBNL, reducing time required to get interesting analysis from DBNL on agent behavior. You can learn more about this integration here (will need to install the NeMo Agent Toolkit off of the develop branch): &lt;a href=&quot;https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/develop/docs/source/workflows/observe/observe-workflow-with-dbnl.md&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/develop/docs/source/run-workflows/observe/observe-workflow-with-dbnl.md&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;steps-to-use-the-dbnl-and-nvidia-nemo-agent-toolkit-integration&quot;&gt;&lt;a href=&quot;#steps-to-use-the-dbnl-and-nvidia-nemo-agent-toolkit-integration&quot;&gt;Steps to use the DBNL and NVIDIA NeMo Agent Toolkit integration&lt;/a&gt;&lt;/h2&gt;
&lt;h3 id=&quot;step-1-install-dbnl&quot;&gt;&lt;a href=&quot;#step-1-install-dbnl&quot;&gt;Step 1: Install DBNL&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Visit &lt;a href=&quot;https://docs.dbnl.com/get-started/quickstart&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://docs.dbnl.com/get-started/quickstart&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; to install DBNL.&lt;/p&gt;
&lt;h3 id=&quot;step-2-create-a-project&quot;&gt;&lt;a href=&quot;#step-2-create-a-project&quot;&gt;Step 2: Create a project&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Create a new Trace Ingestion project in DBNL. To create a new project in DBNL:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Navigate to your DBNL deployment (e.g. &lt;a href=&quot;http://localhost:8080/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;http://localhost:8080/&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Go to Projects &amp;gt; + New Project&lt;/li&gt;
&lt;li&gt;Name your project nat-calculator&lt;/li&gt;
&lt;li&gt;Add a LLM connection to your project&lt;/li&gt;
&lt;li&gt;Select Trace Ingestion as the project Data Source&lt;/li&gt;
&lt;li&gt;Click on Generate API Token and note down the generated API Token&lt;/li&gt;
&lt;li&gt;Note down the Project ID for the project&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&quot;step-3-configure-your-environment&quot;&gt;&lt;a href=&quot;#step-3-configure-your-environment&quot;&gt;Step 3: Configure your environment&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Set the following environment variables in your terminal:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# DBNL_API_URL should point to your deployment API URL (e.g. http://localhost:8080/api)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export DBNL_API_URL=&amp;lt;your_api_url&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export DBNL_API_TOKEN=&amp;lt;your_api_token&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export DBNL_PROJECT_ID=&amp;lt;your_project_id&amp;gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;step-4-install-the-nemo-agent-toolkit-opentelemetry-subpackages&quot;&gt;&lt;a href=&quot;#step-4-install-the-nemo-agent-toolkit-opentelemetry-subpackages&quot;&gt;Step 4: Install the NeMo Agent Toolkit OpenTelemetry subpackages&lt;/a&gt;&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Install specific telemetry extras required for DBNL&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;uv pip install -e &amp;#39;.[opentelemetry]&amp;#39;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;step-5-modify-nemo-agent-toolkit-workflow-configuration&quot;&gt;&lt;a href=&quot;#step-5-modify-nemo-agent-toolkit-workflow-configuration&quot;&gt;Step 5: Modify NeMo Agent Toolkit workflow configuration&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Update your workflow configuration file to include the telemetry settings.&lt;/p&gt;
&lt;p&gt;Example configuration:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;general:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  telemetry:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    tracing:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      dbnl:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        _type: dbnl&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;step-6-run-the-workflow&quot;&gt;&lt;a href=&quot;#step-6-run-the-workflow&quot;&gt;Step 6: Run the workflow&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;From the root directory of the NeMo Agent Toolkit library, install dependencies and run the pre-configured simple_calculator_observability example.&lt;/p&gt;
&lt;p&gt;Example:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Install the workflow and plugins&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;uv pip install -e examples/observability/simple_calculator_observability/&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Run the workflow with DBNL telemetry settings&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Note: you may have to update configuration settings based on your DBNL deployment&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nat run --config_file examples/observability/simple_calculator_observability/configs/config-dbnl.yml --input &amp;quot;What is 1*2?&amp;quot;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;As the workflow runs, telemetry data will start showing up in DBNL.&lt;/p&gt;
&lt;h3 id=&quot;step-7-analyze-traces-in-dbnl&quot;&gt;&lt;a href=&quot;#step-7-analyze-traces-in-dbnl&quot;&gt;Step 7: Analyze traces in DBNL&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;To analyze traces in DBNL:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Navigate to your DBNL deployment (e.g. &lt;a href=&quot;http://localhost:8080/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;http://localhost:8080/&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Go to Projects &amp;gt; nat-calculator&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For additional help, see the &lt;a href=&quot;https://docs.dbnl.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;DBNL docs&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Here is an example of tool sequence flow analysis in Distributional:&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/integrating-dbnl-with-the-nvidia-nemo-agent-toolkit/6931e84c27ea89d09016bc27_image2.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;And here is an example insight and evidence you should expect to see:&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/integrating-dbnl-with-the-nvidia-nemo-agent-toolkit/6931e863a948887ed921cfd7_image4.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;next&quot;&gt;&lt;a href=&quot;#next&quot;&gt;Next&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;With this setup, it should be easy for you to start publishing and analyzing traces from your NeMo Agent Toolkit agent in DBNL. If you are interested in other integrations, reach out at &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;contact@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; to request one and we’ll prioritize it.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Adapting product analytics to the AI era</title><link>https://distributional.com/blog/adapting-product-analytics-to-the-ai-era</link><guid isPermaLink="true">https://distributional.com/blog/adapting-product-analytics-to-the-ai-era</guid><description>Scott Clark&apos;s pivot-to-analytics thesis: production AI as a black box, behavior as the unit of analysis, and production-led development. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Tue, 04 Nov 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;h2 id=&quot;production-ai-is-a-black-box&quot;&gt;&lt;a href=&quot;#production-ai-is-a-black-box&quot;&gt;Production AI is a black box&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;“We are flying blind,” said the data platform leader at a large re-insurance provider. “We log traces, but lack a reliable way to analyze them—so they just sit there. I’m under pressure to deliver returns from this massive AI investment, and it starts with a better understanding of which AI use cases are most valuable. This insight is hidden in these logs somewhere.”&lt;/p&gt;
&lt;p&gt;A year ago, her team couldn’t get governance approvals to use LLMs. Once they overcame this hurdle, she released a chat experience internally. Then she added RAG to this solution, expanding the types of questions her teams could ask and the relevance of answers they would expect back. More recently, she added an agentic router. Now she has thousands of users with tens of thousands of requests per day. But her AI product is a black box and she is missing the tools to navigate it.&lt;/p&gt;
&lt;p&gt;She is not alone. I’ve met with dozens of companies that are asking similar questions, and finding it harder than expected to get answers. Since this conversation, we’ve been working with the data platform leader and folks like her at other companies to solve this problem.&lt;/p&gt;
&lt;h2 id=&quot;ai-product-behavior&quot;&gt;&lt;a href=&quot;#ai-product-behavior&quot;&gt;AI product behavior&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;If you have any of these problems, or if you struggle with your own unique variant, you may be inclined to pick up a tried and true DevOps solution, like traditional monitoring or product analytics tools. But AI applications come with their own unique attributes that break these products.&lt;/p&gt;
&lt;p&gt;Unstructured or semi-structured data makes it hard to statistically analyze usage of these applications. When you are given piles of data, it is non-trivial to unpack what matters and what you can ignore. The high degree of non-determinism implicit in these models that makes them so powerful is also what makes them hard to evaluate. This non-determinism means that a lot of issues, insights, or information can be hidden in a summary metric. And both of these problems are compounded by the complexity of the typical AI application. Context sources, retrieval mechanisms, re-rankers, tools, tasks, and routers are all examples of components outside the AI model that are constantly changing. The need to assess and interpret traces rather than specific data points makes this problem even harder.&lt;/p&gt;
&lt;p&gt;Unstructured data, non-determinism, and multi-component complexity are all problems that have been dealt with in some form or another, but the way that these problems collide with AI applications requires they be treated with a purpose-built AI product analytics solution.&lt;/p&gt;
&lt;p&gt;The key to peering into the production AI black box is understanding AI &lt;em&gt;behavior&lt;/em&gt;—the interplay and correlations between users, context, tools, models, and metrics.&lt;/p&gt;
&lt;h2 id=&quot;our-workflow-gives-you-purchase-to-climb-this-mountain&quot;&gt;&lt;a href=&quot;#our-workflow-gives-you-purchase-to-climb-this-mountain&quot;&gt;Our workflow gives you purchase to climb this mountain&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional’s product is designed to help you understand your product as it evolves. Our platform automatically learns the graph of user experiences with your product, classifies these experiences, and then serves daily insights on them. As you take action on these insights, Distributional tailors future insights to these preferences. It gives you the initial toehold, and the more you use it the higher you climb the mountain.&lt;/p&gt;
&lt;p&gt;Distributional enables this experience with three components: pipeline, workflow, and platform.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/adapting-product-analytics-to-the-ai-era/68f7b134322cf88f48804994_Distributional-diagram-alone-compress.gif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Learn more: &lt;a href=&quot;https://docs.dbnl.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://docs.dbnl.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/p&gt;
&lt;h3 id=&quot;platform&quot;&gt;&lt;a href=&quot;#platform&quot;&gt;Platform&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Our goal is to make it as easy as possible to get this pipeline and workflow running in your environment. The platform is architected to fit your existing stack. You can install our full service for free. It runs on your Kubernetes cluster, and no data ever leaves your system. We have built an initial set of data and model connectors, and are happy to extend these primitives to make the connection seamless. The goal is for our product to fit your stack, not for you to redesign your stack to fit our product.&lt;/p&gt;
&lt;p&gt;Learn more: &lt;a href=&quot;https://docs.dbnl.com/platform/deployment&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://docs.dbnl.com/platform/deployment&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/p&gt;
&lt;h3 id=&quot;pipeline&quot;&gt;&lt;a href=&quot;#pipeline&quot;&gt;Pipeline&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Once installed in your environment, Distributional automatically processes your production AI logs to make sense of behavior. This happens in a three step process. First, our product enriches your logs with metrics, statistics, attributes, evals, and LLM as judge metrics to capture the status of AI product behavior and help you understand canonical usage patterns. Second, Distributional analyzes these enriched production logs to uncover behavioral signals with unsupervised clustering, topic modeling, anomaly detection, and data drift assessment. This analysis generates a complete graph of typical usage, organizes usage through classification by highest propensity topics, and produces a set of insights on compelling recent usage patterns or issues. Finally, our product publishes these behavioral signals to a dashboard, your daily report, and session analysis with notifications on significant deviations from historical behavior. As you engage with these published results, our pipeline reinforces these preferences in future analysis.&lt;/p&gt;
&lt;p&gt;Learn more: &lt;a href=&quot;https://docs.dbnl.com/configuration/data-pipeline&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://docs.dbnl.com/configuration/data-pipeline&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;workflow&quot;&gt;&lt;a href=&quot;#workflow&quot;&gt;Workflow&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;This pipeline feeds a workflow designed to help you understand your AI products and complete your AI product feedback loop. First, Distributional produces fresh daily insights that update baseline understanding of usage patterns and alert you to any significant changes, shifts, or deviations. Second, our product links these daily insights to evidence—specific logs, sessions, and charts—so you can perform rapid root cause analysis. As you either make changes to your product, adjust thresholds for existing metrics, or decide to add new metrics, Distributional becomes your source of record for these changes so you can track how the behavior of your AI product evolves over time.&lt;/p&gt;
&lt;p&gt;Learn more: &lt;a href=&quot;https://docs.dbnl.com/workflow/adaptive-analytics-workflow&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;https://docs.dbnl.com/workflow/adaptive-analytics-workflow&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;production-led-ai-development&quot;&gt;&lt;a href=&quot;#production-led-ai-development&quot;&gt;Production-led AI development&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This approach takes some of the challenges of AI applications and turns them into an advantage—insights that your team can use to improve these applications over time.&lt;/p&gt;
&lt;p&gt;It also implies that you should spend as little time upfront on evals as possible, and instead get your AI application in production as soon as you feel comfortable. Then perform robust analysis of production data to guide how you are developing your AI application. Unsupervised techniques are well designed for giving you these insights “for free” on production data.&lt;/p&gt;
&lt;p&gt;Ultimately, this approach brings a few benefits. You’ll ship AI applications faster, avoid degradations as you scale, evolve your applications with your users over time, and be more confident in AI performance by aligning its behavior to your goals.&lt;/p&gt;
&lt;p&gt;Ready to go? &lt;a href=&quot;https://docs.dbnl.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Install&lt;/a&gt; our product free and get started today. And I’m always happy to chat about it with you, so find me on &lt;a href=&quot;https://www.linkedin.com/in/nick-payton/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;LinkedIn&lt;/a&gt; or email &lt;a href=&quot;mailto:nick-dbnl@distributional.com&quot;&gt;nick-dbnl@distributional.com&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>The Distributional Workflow: Define, Detect, Understand, and Improve</title><link>https://distributional.com/blog/the-distributional-workflow-define-detect-understand-and-improve</link><guid isPermaLink="true">https://distributional.com/blog/the-distributional-workflow-define-detect-understand-and-improve</guid><description>The Define, Detect, Understand, Improve workflow walked end to end on a customer-support agent. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Tue, 26 Aug 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Distributional is an adaptive testing solution for enterprise AI applications, built to help teams define, detect, understand, and improve upon an application’s desired behavior. To enable this, Distributional’s user experience was designed to translate the complexity of GenAI application logs and traces into coherent insights, giving teams rich information to act on. By providing several opportunities to customize the experience, Distributional can be tailored to an individual team or app’s specific needs.&lt;/p&gt;
&lt;p&gt;In this article, we’ll review an example of a customer support agent use case to showcase how the product works, and map it to the four-stage Distributional workflow: Define, Detect, Understand, and Improve. To learn more, download the paper on &lt;a href=&quot;http://www.distributional.com/papers/distributionals-user-experience&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Distributional’s User Experience&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;
&lt;h2 id=&quot;the-distributional-workflow&quot;&gt;&lt;a href=&quot;#the-distributional-workflow&quot;&gt;The Distributional workflow&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;There are four steps in the Distributional workflow:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Define&lt;/strong&gt; the behavior or status of your AI application&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Detect&lt;/strong&gt; changes in this behavior over time&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Understand&lt;/strong&gt; what caused this change and whether you care&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Improve&lt;/strong&gt; measurement of AI application behavior to align it with your business goals&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The expectation is that AI product and engineering teams iterate through these steps on a regular basis as their AI application evolves in production. As they iterate through these steps, they refine and improve their definition of AI application behavior, leading to even richer insights in the future. This virtuous cycle of continuous AI product improvement is designed to help teams keep up with fast-paced AI innovations and changing user behavior.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4c6e7a5af08ea9b27253b_AD_4nXebWe9yAZ1zdKDA7REdcFIbzY39ujezjvmvchaWT5sYtWydspXpLdaGF9VTvlyYiPy9QGL6Y9t10JGEA3vgFLVB_6_l9QcPARHVSUW-Th5JhfLS5sj2J13UkKCMHUPFILwDTo-N.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;define&quot;&gt;&lt;a href=&quot;#define&quot;&gt;Define&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional’s workflow begins by walking a user through the process of defining their application’s behavior.&lt;/p&gt;
&lt;p&gt;The Distributional platform automatically quantifies AI application behavior off of unstructured data with the Eval Module, which is designed to significantly increase the number of metrics being tracked. You are also able to add your own metrics to this, though it’s not required, so that the platform can provide a more robust quantitative view of the application’s behavior.&lt;/p&gt;
&lt;p&gt;Similarly, the Distributional platform automatically assesses similarity using the &lt;a href=&quot;https://distributional.com/blog/distributional-simplifies-adaptive-testing-with-similarity-index-and-key-insights&quot;&gt;Similarity Index&lt;/a&gt; (Sim Index), which is derived from statistical tests comparing these metrics across experiment and baseline runs and presented as an aggregate value that represents how much behavior has shifted. Distributional also includes a library of a wide variety of statistical tests so you can easily add your own tests or adjust thresholds—either programmatically or with a single click.&lt;/p&gt;
&lt;p&gt;Together, these items create your app’s Behavioral Fingerprint. Behavioral because it represents information on AI application status beyond performance, inclusive of a variety of attributes that may be correlated in interesting ways. And Fingerprint because the combination of distributions of metrics that define their status will be unique to each AI application.&lt;/p&gt;
&lt;p&gt;The goal of the rest of this workflow is to consistently use information from the other steps to refine the definition of AI app status so it is always consistent with the current state of the AI application as it evolves over time.&lt;/p&gt;
&lt;h3 id=&quot;agent-use-case&quot;&gt;&lt;a href=&quot;#agent-use-case&quot;&gt;Agent use case&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;In our use case, the agent receives an input, extracts entities and sentiment from the metadata, infers the task, and then defines routing priority. This routing priority step then dictates the response to the user and the final routing output. (If you are more of a visual learner, you can view the &lt;a href=&quot;https://youtu.be/F2e9_UxyrVs?si=evCNOt9F8S-bZaO9&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;demo video&lt;/a&gt; instead.)&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4e6bb6212ae7bf94aea5a_image8.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;For the response to the user, we include nine metrics, most of which are intended to contextualize whether the LLM output is consistent with expectations. These metrics include automated readability index, custom friendliness metric, Flesch Kincaid grade, LLM grammatical accuracy, LLM reading complexity, LLM sentiment assessment, LLM text toxicity, token count, and word count. This mixes a variety of types of metrics, including statistical metrics, custom user defined metrics, LLM as judge metrics, business metrics, and properties of the underlying data.&lt;/p&gt;
&lt;p&gt;For the intermediate steps, we also have four metrics in this example. Although you can catch most issues with assessment of response to user metrics, these intermediate steps metrics make root cause analysis much more streamlined and guided. These metrics include sentiment confidence score, sentiment polarity, routing priority, and task inferred. Below is an example of a variety of these and the representation of summary statistics on each distribution in our product.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4c73d46da8826d2a7b4cf_AD_4nXdmHtR7GzB_A8jmZa9IF3gS7tgRFxLu-wEcf0GyrtMTRmHnuq1ks4GTs-Onk_q83nFTZMmVBRd8yKL90uV1UwA7LjMEFNYktW3NF7rXPN7wg3HOstGvWwbRNn_rQNI7cw2ZNGajGg.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;detect&quot;&gt;&lt;a href=&quot;#detect&quot;&gt;Detect&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;In Distributional, detection is both automated and configurable. The platform automatically applies the Sim Index to detect any change over two points in time for the application as a whole and at the  component, column, and result levels so you have an intuitive way to detect and identify where there has been a shift. While the Sim Index comes with a pre-configured threshold, it is easy to adjust so you can minimize false positives or false negatives.&lt;/p&gt;
&lt;p&gt;Along with these thresholds, you can configure notifications and alerting on Similarity Indexes or other tests directly within the platform with a few clicks, and customize the notifications so they are aligned with the severity of the change. For example, drastic shifts around tokens or toxicity may warrant a PagerDuty alert, while subtle shifts on output relevance may only require a Slack notification. Learn more about notifications &lt;a href=&quot;https://docs.dbnl.com/using-distributional/notifications&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;here&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;A key tenet of Distributional’s workflow is that detecting shifts will yield new insights on the behavior of your AI application, and that you’ll want to use this insight to further improve the definition of behavior by creating new tests. To support this, it’s easy to set new thresholds or make new assertions in Distributional, either within the dashboard or programmatically through the SDK.&lt;/p&gt;
&lt;h3 id=&quot;agent-use-case-1&quot;&gt;&lt;a href=&quot;#agent-use-case-1&quot;&gt;Agent use case&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;A week into the production application, Distributional identifies a significant change in behavior, as represented by a run-over-run Similarity Index of 33 and is notified which tests failed—in this case, all of them.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4e6f8f60fd1a27c685e3f_image4.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;understand&quot;&gt;&lt;a href=&quot;#understand&quot;&gt;Understand&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;For many AI teams, trying to root cause and debug issues once detected can take hours. Distributional is designed to give you relevant insights on any change as quickly and intuitively as possible so you can do rapid root cause analysis and take appropriate action. To simplify this process, Distributional helps you quickly answer these three straightforward questions:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Was there a change?&lt;/li&gt;
&lt;li&gt;Where was the change?&lt;/li&gt;
&lt;li&gt;Do I care about the change?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;To answer these, you have a few tools available within the platform:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The &lt;a href=&quot;https://docs.dbnl.com/using-distributional/tests/reviewing-tests&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;&lt;strong&gt;Similarity Report&lt;/strong&gt;&lt;/a&gt; provides an overview of all shifts in Sim Indexes at the app and column level, along with a human readable description of the shift. It also enables guided investigation into column specific changes.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.dbnl.com/using-distributional/tests/reviewing-tests#similarity-insights&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;&lt;strong&gt;Similarity Insights&lt;/strong&gt;&lt;/a&gt; are automatically produced with every Sim Index and provide human readable insights that represent the most probable causes of the change. These drastically accelerate the ability to root cause any changes.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.dbnl.com/using-distributional/tests/reviewing-tests/notable-results&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;&lt;strong&gt;Notable Results&lt;/strong&gt;&lt;/a&gt; are specific rows of data that are automatically surfaced for deeper, line-by-line review during investigation.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These features are paired with correlations, comparative analysis, and an array of visualizations designed to guide your understanding on the change. Additionally, you also have the full history of data for all components in your AI system (with admin configurable access controls and permissions), so you can go as deep as you want, and share the information with other teams for collaboration on any action that you take to resolve the issue or improve the AI app.&lt;/p&gt;
&lt;h3 id=&quot;agent-use-case-2&quot;&gt;&lt;a href=&quot;#agent-use-case-2&quot;&gt;Agent use case&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;First, Distributional provides Similarity Insights that immediately offer the user guidance on what most likely caused the change. In this case, routing priority is served up as the metric to analyze first.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4e72665d87e569eb9895b_image4b.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Second, Distributional offers a Similarity Report to see these Similarity Insights in context of everything that has changed to get a full picture of how the metrics have varied over time. Again, this Report suggests looking at routing priority first, but also contextualizes this by denoting that task inferred also experienced a significant shift as well.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4c783dccbf1d3ce0b0a2c_AD_4nXedtWrAtBKTyOGyKfrK2P2SLxMHpKzkztv16gWZ1gBzH8dm9kgUSc3oaNOnHo929qk6FHD_aV7h3bcvbC21WYmFaVW1agC5silq2hUwNetPvaDNgI1CKKgIR2OzFBrKl--GYvWJ.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Distributional then offers up comparative analysis and other visualizations of these distributions to help you understand what has happened. In this case, it is obvious that the routing priority has shifted to have a much higher propensity of &lt;strong&gt;high&lt;/strong&gt; routing priority inputs.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4e769bb930a85e2f5a2a2_image2.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Finally, Distributional then provides specific Notable Results—rows of data to be analyzed that most contextualize this change. This shifts data review from random sampling guesswork to targeted, precise review of specific prompt-response pairs that need attention. In this case, it is quickly obvious that the same responses are generating wildly different routing priority.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4e78066c55cad0db963f7_image2b.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;improve&quot;&gt;&lt;a href=&quot;#improve&quot;&gt;Improve&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;One of the most powerful parts of this workflow is the ability to continuously improve your AI application over time.&lt;/p&gt;
&lt;p&gt;This can be in the form of a single click to &lt;a href=&quot;https://docs.dbnl.com/using-distributional/tests/creating-tests&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;create a new test&lt;/a&gt; on a metric that you discovered had a certain behavior you want to track. Or it could be selecting rows of data the platform surfaces through Notable Results and building these into your golden dataset for future development. In other instances, it could be providing a dashboard report to governance teams for approval to scale to new segments or users after certifying behavior through the tests being run. Finally, teams will often rely on Distributional for continuous app development (such as refactors, upgrades, or otherwise) to ensure that any development won’t fail this collection of tests that define “desired” AI application behavior.&lt;/p&gt;
&lt;p&gt;In all cases, the Distributional platform is designed to flexibly fit into whichever version of this workflow you have put in place for your team.&lt;/p&gt;
&lt;h3 id=&quot;agent-use-case-3&quot;&gt;&lt;a href=&quot;#agent-use-case-3&quot;&gt;Agent use case&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;With the direction that Distributional offers, this team discovered a change in the prompt template that drove this shift in behavior, and were able to resolve it. They re-ran tests to confirm this fix passed their set of statistical tests that defined good, performative AI application behavior.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4c7aaf438e409660fbe01_AD_4nXfMeONLyoyNIFBc4Mpeifb8PzIMZ5FVPm8nZ4VSgnn2gRElIiHhS3nRA4toD7nEIQpGp8oFykm-GmEbkkLltn8zFo249EOmh-eGgwC8IAu-1LFLC6tXSbW7fvoQg1mOFXg0ZraNLA.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Similarly, they were able to add three new tests guided by the Distributional platform that represented a more rigorous threshold and definition of behavior they wanted to track. In this case, this included new discrepancy tests for sentiment polarity, routing priority, and task inferred so they could be notified regarding deviations in the future.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-distributional-workflow-define-detect-understand-and-improve/68a4c7c48785bcc2796cff1b_AD_4nXclgDjW4PgamBqkaeR4HutLdeYyS91ZqcWu_0JLT6snnGylQAiihBe0NFMiYDf4K0W67bFd27tLnh509ue2hih9Cm6LNxsXURevxIVjQosRbT2tFmuyn9QLJ8Zom_KzllLVPYCh2A.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h2 id=&quot;result&quot;&gt;&lt;a href=&quot;#result&quot;&gt;Result&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This use case showcases the value propositions in the user experience. First is the value of Distributional’s automation in streamlining the process of defining AI application behavior and writing tests to check this behavior over time. Second, the Distributional platform provides intelligent recommendations for what is causing any change that simplifies the root cause analysis process. Finally, the platform serves as a system of record for all tests and analysis so teams can reproduce what happened and share it with third parties.&lt;/p&gt;
&lt;p&gt;While this is just one use case, this is meant to give you a sense of the range of areas where you’d gain insights with Distributional. No matter what your app looks like, Distributional’s goal is to give you a consistent way to assess and analyze each of these components.&lt;/p&gt;
&lt;h2 id=&quot;learn-more&quot;&gt;&lt;a href=&quot;#learn-more&quot;&gt;Learn more&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional aims to provide an intuitive experience for understanding what, where, why, and how your AI applications are changing over time, so you can use this information to guide continuous development. For a deeper dive into the Distributional workflow, check out the downloadable paper on &lt;a href=&quot;http://www.distributional.com/papers/distributionals-user-experience&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Distributional’s User Experience&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Scale, Metrics, Tests, and Results in Distributional</title><link>https://distributional.com/blog/scale-metrics-tests-and-results-in-distributional</link><guid isPermaLink="true">https://distributional.com/blog/scale-metrics-tests-and-results-in-distributional</guid><description>How scale, metrics, tests, and results compose Distributional&apos;s analysis model for translating unstructured LLM data into testing. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Mon, 18 Aug 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Generative AI is a platform shift that promises to fundamentally change software. But what makes GenAI so powerful—including its non-stationary and non-deterministic nature—also makes its behavior that much more complex to understand and manage, especially at scale.&lt;/p&gt;
&lt;p&gt;Today’s GenAI apps often contain multiple moving parts that each have their own data and non-determinism. Attributes like endpoint consistency, endpoint refactor consistency, retrieval efficacy, embedding pipeline efficacy, and components related to agents require that AI teams rely on adaptive tools to gain clarity on where a shift in behavior has occurred and how to address it. A more comprehensive, adaptive way of understanding this is needed.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.distributional.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Distributional&lt;/a&gt; is an adaptive testing solution for enterprise AI applications, built to help teams define, detect, understand, and improve upon an application’s desired behavior. In this article, we’ll review the four key concepts within Distributional to help you gain an understanding of how the platform works, including: Scale, Metrics, Tests, and Results. To learn more, download the paper on &lt;a href=&quot;http://www.distributional.com/papers/distributionals-user-experience&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Distributional’s User Experience&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; or visit our &lt;a href=&quot;https://docs.dbnl.com/learning-about-distributional/distributional-concepts&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;docs&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;
&lt;h2 id=&quot;distributional-concepts&quot;&gt;&lt;a href=&quot;#distributional-concepts&quot;&gt;Distributional concepts&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional is designed to handle scale and complexity. Our architecture automatically scales batch analysis of underlying LLM data, and our tests are designed to translate unstructured data into quantitative analysis that scales with usage of your AI applications. Both are designed to handle the full complexity of your application, testing and analyzing data for every component to make root cause analysis faster and more effective.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/scale-metrics-tests-and-results-in-distributional/689e67b91fe01ae894705bc8_image1.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h3 id=&quot;scale&quot;&gt;&lt;a href=&quot;#scale&quot;&gt;Scale&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;We designed this workflow to help you &lt;a href=&quot;https://docs.dbnl.com/learning-about-distributional/the-flow-of-data&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;scale testing and analysis&lt;/a&gt; of Generative AI applications in production. It is &lt;strong&gt;not&lt;/strong&gt; a workflow designed for rapid iteration and prototyping in development—there are plenty of open source libraries with this workflow.&lt;/p&gt;
&lt;p&gt;You should start using Distributional when you already feel okay about the performance of your LLM application, using our solution to robustly define this baseline behavior and then continuously assessing drift from this baseline expectation. Our goal is to give you insights to production usage and performance of your AI application as part of a continuous testing, analysis, and development process.&lt;/p&gt;
&lt;h3 id=&quot;metrics&quot;&gt;&lt;a href=&quot;#metrics&quot;&gt;Metrics&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Distributional’s approach to &lt;a href=&quot;https://docs.dbnl.com/using-distributional/metrics&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;metrics&lt;/a&gt; is the more the merrier. When you point our platform at your GenAI application production logs, our &lt;a href=&quot;https://docs.dbnl.com/reference/python-sdk/eval-module&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Eval Module&lt;/a&gt; automatically generates dozens of statistical and LLM-as-a-Judge metrics to transform this unstructured data into testable quantities.&lt;/p&gt;
&lt;p&gt;Our goal with this library is to expand the number of quantities you are automatically testing so you have a broader, more robust definition of overall AI app behavior. By expanding the number of metrics, the idea is that a wide variety of weak estimators will give you a more durable assessment of change in performance over time, rather than a few narrowly defined evals. We do, however, recommend that you also upload any evals you have constructed, so you can observe these in the context of the wider variety of metrics the Distributional platform provides.&lt;/p&gt;
&lt;p&gt;The benefit of this approach is that you can discover correlations across these metrics that may yield interesting insight on AI app behavior that you can use to continuously evolve your application. These metrics and the actual underlying unstructured data from every component form columns in our platform.&lt;/p&gt;
&lt;h3 id=&quot;tests&quot;&gt;&lt;a href=&quot;#tests&quot;&gt;Tests&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;After computing metrics from runs, the Distributional platform automatically applies statistical &lt;a href=&quot;https://docs.dbnl.com/using-distributional/tests&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;tests&lt;/a&gt; to compare the distribution of these metrics in the baseline run versus the experiment run. These tests are usually designed to assess similarity over time where the baseline run is day zero or day t-1 and the experiment run is today. They can also be designed to compare model version A as experiment versus model version B as the baseline. The goal of these statistical tests is to assert a threshold level of consistency run-over-run.&lt;/p&gt;
&lt;p&gt;Our product contains over 60 statistical tests that can be applied programmatically or with a click of a button. But to simplify the initial user experience, we start with a single test for similarity using our &lt;a href=&quot;https://docs.dbnl.com/using-distributional/tests/reviewing-tests/what-is-a-similarity-index&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Similarity Index&lt;/a&gt; (Sim Index). This approach means that you will start to get insights on shifts in AI app behavior without having to do any work designing your own tests or parsing each of these dozens of statistical tests. Instead, as you learn how the Sim Index is changing over time, this is an opportunity for you to apply any one of these dozens of statistical tests with a single click of a button.&lt;/p&gt;
&lt;p&gt;The advantage of this approach is that it also scales. You can observe our Sim Index on every component of a complex agent or RAG system to understand which component is driving change. You can also view our Sim Index at the Run level across every application to understand which may need more attention.&lt;/p&gt;
&lt;h3 id=&quot;results&quot;&gt;&lt;a href=&quot;#results&quot;&gt;Results&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Our &lt;a href=&quot;https://docs.dbnl.com/using-distributional/tests/reviewing-tests&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Similarity Report&lt;/a&gt; provides an overview of the Similarity Index for all metrics, giving you a quick view into what has changed so you can assess whether there is anything of note. The idea with this Report is that it should enable AI product teams to get on the same page, and also give that team a view to share with governance or leadership teams that need information on what has been tested but may not have complete AI application context.&lt;/p&gt;
&lt;p&gt;Bundled with this Report are &lt;a href=&quot;https://docs.dbnl.com/using-distributional/tests/reviewing-tests#similarity-insights&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Similarity Insights&lt;/a&gt;, which are our platform’s assessment of which changes are most impactful, and where you should focus your attention when starting to diagnose the issue. These Similarity Insights translate statistical tests into human readable insights on the drivers of this change.&lt;/p&gt;
&lt;p&gt;Finally, we also automatically offer &lt;a href=&quot;https://docs.dbnl.com/using-distributional/tests/reviewing-tests/notable-results&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Notable Results&lt;/a&gt; that feed you specific rows of data that are most important to prioritize when performing root cause analysis. This eliminates the need to sift through rows of data or randomly sample it to try to get intuition on the user experience. Instead, you get a direct signal on which traces are most important to review.&lt;/p&gt;
&lt;p&gt;These insights—Similarity Index, Similarity Insights, and Notable Results—are designed to provide an intuitive workflow to quickly understand what has changed, what is driving it, and whether you care.&lt;/p&gt;
&lt;h2 id=&quot;learn-more&quot;&gt;&lt;a href=&quot;#learn-more&quot;&gt;Learn more&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;For a deeper dive into Distributional concepts and to learn more about how to use Distributional, check out the downloadable paper on &lt;a href=&quot;http://www.distributional.com/papers/distributionals-user-experience&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Distributional’s User Experience&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>How Distributional tests for consistent app behavior in production</title><link>https://distributional.com/blog/how-distributional-tests-for-consistent-app-behavior-in-production</link><guid isPermaLink="true">https://distributional.com/blog/how-distributional-tests-for-consistent-app-behavior-in-production</guid><description>Prompts and responses as random variables: the hypothesis-testing framing behind testing AI apps for consistent production behavior. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Fri, 15 Aug 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;For teams scaling AI applications in production, adaptive testing is critical for consistency and reliability. But testing these applications is not trivial. Due to the nature of the embedding process and subsequent recovery of text from embedding space, these apps require statistical analysis and testing. Additionally, once an app is running in production, testing automation becomes necessary since teams can no longer manually test for all possible usages and analyze all possible responses.&lt;/p&gt;
&lt;p&gt;Distributional provides an adaptive testing platform that is uniquely designed to address these needs. Distributional’s testing strategy is based on analyzing recent production usage to test for consistency of app behavior as a whole, while also providing mechanisms to both alert users to behavioral deviations and provide interpretable evidence for users to understand what has occurred.&lt;/p&gt;
&lt;p&gt;In this article, we’ll dive into how Distributional tests for consistent app behavior in production.&lt;/p&gt;
&lt;h2 id=&quot;strategy-to-test-for-consistent-app-behavior-in-production&quot;&gt;&lt;a href=&quot;#strategy-to-test-for-consistent-app-behavior-in-production&quot;&gt;Strategy to test for consistent app behavior in production&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;At its core, the Distributional platform uses a testing strategy that embraces the idea that every time an AI app is used, both the prompt (input) and response (output) are random variables drawn from some distribution. The goal is to analyze the behavioral consistency of the app. For example, has the app responded to questions on a specific topic differently today than yesterday?&lt;/p&gt;
&lt;p&gt;The platform then surfaces notable deviations to users in an unsupervised way. By design, it does not pass judgment on whether the responses are correct or not—rather, Distributional gives users the ability to introspect and apply supervision by passing judgment based on their own expertise and their use cases. The platform is able to analyze input/output consistency from an app, as well as consistency across intermediate data generated by the app.&lt;/p&gt;
&lt;p&gt;To start to visualize this testing strategy, capital letters represent a random variable, and lowercase letters represent a realization of that random variable or other deterministic quantity:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;𝐼 – the distribution of possible inputs (prompts) to the app&lt;/li&gt;
&lt;li&gt;𝑖 – an observed set of prompts&lt;/li&gt;
&lt;li&gt;𝑂 – the distribution of possible outputs (responses) from the app&lt;/li&gt;
&lt;li&gt;𝑜 – an observed set of outputs&lt;/li&gt;
&lt;li&gt;𝑡 – time span over which inputs and outputs are considered&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Next, some conditional notation to facilitate analysis in Distributional:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;I | t1&lt;/em&gt; – the possible inputs over time span t1&lt;/li&gt;
&lt;li&gt;&lt;em&gt;I, O | t1&lt;/em&gt; – the possible inputs and outputs over time span t1&lt;/li&gt;
&lt;li&gt;&lt;em&gt;O | t1, I&lt;/em&gt; – the possible outputs over time span t1, given the inputs I&lt;/li&gt;
&lt;li&gt;&lt;em&gt;tb&lt;/em&gt; – a baseline time period against which recent usage is compared&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Distributionalʼs core testing functionality studies the following hypothesis test:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;H0&lt;/em&gt; – &lt;em&gt;I, O | te&lt;/em&gt; and &lt;em&gt;I, O | tb&lt;/em&gt; are the same distribution&lt;/li&gt;
&lt;li&gt;&lt;em&gt;H1&lt;/em&gt; – They are not the same distribution&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In general, &lt;em&gt;te&lt;/em&gt; is a recent time period and &lt;em&gt;tb&lt;/em&gt; is a fixed time window from the past. For example, comparing usage from the past 24 hours to usage from last Monday.&lt;/p&gt;
&lt;p&gt;Distributional then helps users understand whether they should reject H0 by presenting relevant evidence of behavioral deviations based on logged app usage between &lt;em&gt;tb&lt;/em&gt; and &lt;em&gt;te&lt;/em&gt;. If the user finds this evidence compelling, notifications can be created to identify such behavior in the future.&lt;/p&gt;
&lt;p&gt;This evidence can also be used to consider alternate, more targeted, null hypotheses. These could be:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Ho&lt;/em&gt;: &lt;em&gt;I | te&lt;/em&gt; and &lt;em&gt;I | tb&lt;/em&gt; are the same distribution&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;which asks only whether the inputs to the app have changed or&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Ho&lt;/em&gt;: &lt;em&gt;O | I, te&lt;/em&gt; and &lt;em&gt;O | I, tb&lt;/em&gt; are the same distribution&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;which asks only whether the outputs have changed given the input distribution.&lt;/p&gt;
&lt;h2 id=&quot;background-on-distributional-consistency&quot;&gt;&lt;a href=&quot;#background-on-distributional-consistency&quot;&gt;Background on distributional consistency&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional consistency could be naturally analyzed under special circumstances. For example, the classic method of the t-test would be a logical strategy for considering consistency of normally distributed data. If we were studying only the inputs I|t and they consisted of only a single numerical value that was normally distributed, then we could analyze whether &lt;em&gt;E [I | tb] ≠ E [I | te]&lt;/em&gt; with a t-test. But in the text-first world of generative AI, it is unlikely that such parametrized analysis will ever be sufficient.&lt;/p&gt;
&lt;p&gt;A nonparametric analysis of distributional similarity has been well addressed for a single continuous or discrete random variable by tools such as the Kolmogorv-Smirnov statistic or Chi-squared statistic, respectively. Tools such as the Kullbeck Liebler (KL) Divergence provide a strategy to measure dissimilarity between random variables when the distribution of those random variables is known.&lt;/p&gt;
&lt;p&gt;However, these tools alone are generally insufficient to analyze the consistency of observed app behavior since:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;the nature of text is neither numerical nor categorical;&lt;/li&gt;
&lt;li&gt;we desire to study multiple variables representing characteristic behavior of the text simultaneously; and&lt;/li&gt;
&lt;li&gt;we are unable to proactively sample data (for, e,g., the KL divergence) given fixed historical logs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Nevertheless, Distributional does incorporate these quantities as facets for helping to define how severely an app’s recent behavior has deviated from previously observed behavior.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/how-distributional-tests-for-consistent-app-behavior-in-production/686d76e5efe950bfa3581b3a_AD_4nXd-LzDw7b9dCAPOOsUc3AYoitecr1eH9hhDWkTFsEvsDXciWXoXRmONUYVJZdUhOS5R139-p2jlps2d0L8igHu19LfolP-oaAI3aqrha793FIcKHlljmNeCowEvxgWbjNKc7G2EKQ.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Evidence of a perceived deviation in behavior may be summarized with a statement such as “Distribution moderately drifted to the left,” but users can dig deeper to interrogate that evidence and judge if this change in behavior is worrisome.&lt;/p&gt;
&lt;h2 id=&quot;incorporating-interpretable-evaluation-metrics&quot;&gt;&lt;a href=&quot;#incorporating-interpretable-evaluation-metrics&quot;&gt;Incorporating interpretable evaluation metrics&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional is designed to empower users to ultimately make the judgment whether a significant or worrisome behavioral deviation has occurred. This means the platform must clearly provide understandable evidence of deviations to users.&lt;/p&gt;
&lt;p&gt;The analysis of &lt;em&gt;H0&lt;/em&gt; is powered by &lt;strong&gt;interpretable evaluation (eval) metrics&lt;/strong&gt;. Distributional provides a set of built-in evals of different structures:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Locally computed classical NLP quantities such as reading level&lt;/li&gt;
&lt;li&gt;LLM-as-judge style quantities leveraging the user’s choice of model&lt;/li&gt;
&lt;li&gt;Hooks for users to create their own LLM-as-judge quantities for submission to Distributional&lt;/li&gt;
&lt;li&gt;RAG-specific quantities to help analyze the behavior of the retrieval process&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Furthermore, any additional quantities that users already compute can be sent to Distributional to be incorporated into the analysis of &lt;em&gt;H0&lt;/em&gt;.&lt;/p&gt;
&lt;h2 id=&quot;download-the-full-tech-paper&quot;&gt;&lt;a href=&quot;#download-the-full-tech-paper&quot;&gt;Download the full tech paper&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional is building the modern enterprise platform for adaptive testing to make AI safe, secure and reliable. As the power of AI applications grows, so does the risk of harm. By taking a proactive, adaptive testing approach with Distributional, AI teams can deploy AI applications with more confidence and catch issues before they cause significant damage in production.&lt;/p&gt;
&lt;p&gt;Learn more about Distributional’s testing strategy by downloading our tech paper on &lt;a href=&quot;https://www.distributional.com/papers/distributionals-approach-to-ai-testing&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Distributional’s Approach to AI Testing&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>How Distributional’s Similarity Index works</title><link>https://distributional.com/blog/how-distributionals-similarity-index-works</link><guid isPermaLink="true">https://distributional.com/blog/how-distributionals-similarity-index-works</guid><description>How the Similarity Index works: computing a 0-100 score per column and aggregating to app level to test behavioral consistency. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Fri, 15 Aug 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Distributionalʼs testing strategy not only needs to analyze the behavioral consistency of an application, but results of this analysis also need to be understandable by users, regardless of the volume of logs, number of metrics, or complexity of the app. This is why Distributional created the &lt;a href=&quot;https://distributional.com/blog/distributional-simplifies-adaptive-testing-with-similarity-index-and-key-insights&quot;&gt;Similarity Index&lt;/a&gt; (Sim Index) as an automatically calculated value—between 0 and 100—which defines the deviation between two time periods, or runs.&lt;/p&gt;
&lt;p&gt;This allows a user to reject H0 in the following hypothesis test:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;H0&lt;/em&gt; – &lt;em&gt;I, O | te&lt;/em&gt; and &lt;em&gt;I, O | tb&lt;/em&gt; are the same distribution&lt;/li&gt;
&lt;li&gt;&lt;em&gt;H1&lt;/em&gt; – They are not the same distribution&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In effect, a lower Sim Index should make the user more likely to reject the null hypothesis. This can be done at the application level, for a grouping of metrics, or for a single metric or eval being logged.&lt;/p&gt;
&lt;p&gt;In this article, we’ll explore Distributional&amp;#39;s strategy behind the Sim Index and how it helps users better understand the behavior of their AI applications.&lt;/p&gt;
&lt;p&gt;(For more information on Distributional’s strategy to test for consistent app behavior in production, including an introduction to the &lt;em&gt;H0&lt;/em&gt; hypothesis, see the &lt;a href=&quot;https://distributional.com/blog/how-distributional-tests-for-consistent-app-behavior-in-production&quot;&gt;previous article&lt;/a&gt; in this series.)&lt;/p&gt;
&lt;h2 id=&quot;computing-the-sim-index&quot;&gt;&lt;a href=&quot;#computing-the-sim-index&quot;&gt;Computing the Sim Index&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The core concept for computing a Sim Index starts with a single column of numerical or categorical data—this could be one of the evals that is computed on the text or a column of non-text data that the user has provided. Let’s refer to this single column as me and mb for the columns from the recent time period and baseline time period, respectively.&lt;/p&gt;
&lt;p&gt;Distributional’s strategy for computing Sim Index on this column, denoted by &lt;em&gt;sim(me,mb)&lt;/em&gt;, is powered by a desire to produce &lt;em&gt;evidence&lt;/em&gt; of dissimilarity between &lt;em&gt;me&lt;/em&gt; and &lt;em&gt;mb&lt;/em&gt;. And then, through this evidence, the user can decide if these columns are, in fact, different in a way that is significant or relevant for their needs.&lt;/p&gt;
&lt;p&gt;To surface such evidence, we assume that &lt;em&gt;H0&lt;/em&gt; is true, i.e., &lt;em&gt;me&lt;/em&gt; and &lt;em&gt;mb&lt;/em&gt; are drawn from the same distribution. We then pool results from both &lt;em&gt;me&lt;/em&gt; and &lt;em&gt;mb&lt;/em&gt; to facilitate an independent and identically distributed (iid) bootstrapping process. This gives an initial sense of what should be happening if the two runs were actually drawn from the same distribution.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/how-distributionals-similarity-index-works/686d780d3dd44192cbc6665d_AD_4nXfldiSgvygj6mF3zc7CUDSGyOvCGYtRRDtvydffT3MqKnjGL4HUcpB-N_2qXfLBCZHwHB3aB3_nPsC4B9ir2eWfZGsuZ-XydLYkW5NU45DDfUhvsk0hQoy1uSTM5013j9degdJK.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Data from today and previously observed data are pooled into a single population from which samples are drawn. These samples, in the box on the far right, provide a sense of what should be observed if todayʼs data showed no deviation from previous data.&lt;/p&gt;
&lt;p&gt;From each of these test pairs, common statistics for both me and mb are computed. These include, for example, the 10th percentile, or the prevalence of the most common categories. This bootstrapping strategy has the benefit of “normalizing” and accounting for the scale of the quantities present to permit all subsequent analyses to take place on a fixed scale. These bootstrapped quantities also provide the desired evidence of behavioral deviation that is surfaced to the user.&lt;/p&gt;
&lt;p&gt;The final Sim Index value for this column is computed through a weighted averaging of the difference of these common statistics as well as nonparametric measures such as the Mann-Whitney U statistic. Future versions of Sim Index may enable users to customize the weighting process to better align with their sense of behavioral deviation, e.g., more heavily weighting tail behavior rather than central tendency of the distribution.&lt;/p&gt;
&lt;h3 id=&quot;sim-index-for-text-columns-and-the-app&quot;&gt;&lt;a href=&quot;#sim-index-for-text-columns-and-the-app&quot;&gt;Sim Index for text columns and the app&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;A Sim Index can also be generated for a text column through combining the Sim Index values for all the quantities derived from that column—for example, the sentiment or toxicity of a given column. That combination is primarily the minimum Sim Index value over the derived metrics. This conservative strategy ensures that any potentially troubling metric deviations are clearly surfaced to users. This also allows users to set thresholds on individual metrics or even statistics of those metrics so they can be notified of any deviations.&lt;/p&gt;
&lt;p&gt;A Sim Index value is also generated at the app-level in the same fashion over all the columns in the app. This provides a helpful signal to users as to whether there have been deviations in behavior overall.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/how-distributionals-similarity-index-works/686d78241f1d64b9771d1fa4_AD_4nXcZpKUrpeXltLIYrdeCkbNcIbF23ysiQuWeQIAqQvq3lzv9Co7pchlEy67S60DB6JFoPI4QTvJ8nuC0Y0BAIxQkKriOl_cP2gQ_5UZtkH2AxE4EwS559ol0UXSEbIcRVIUipmLjrA.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;An app-level Similarity Index, built on column-level Similarity Index values and insights.&lt;/p&gt;
&lt;h2 id=&quot;download-the-tech-paper&quot;&gt;&lt;a href=&quot;#download-the-tech-paper&quot;&gt;Download the tech paper&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional’s approach to AI testing provides a flexible framework that teams can quickly get started with and scale to fit their needs. Through the Sim Index and interpretable metrics, it enables teams with immediate signals on shifts in behaviors, and arms them with the relevant evidence to understand those deviations and pass judgment on whether they matter for their app. Ultimately, providing an automated workflow that AI product teams can use to continuously define, understand, and improve AI application behavior in production.&lt;/p&gt;
&lt;p&gt;Learn more about Distributional’s testing strategy by downloading our tech paper on &lt;a href=&quot;https://www.distributional.com/papers/distributionals-approach-to-ai-testing&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Distributional’s Approach to AI Testing&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>How Distributional balances security with insights</title><link>https://distributional.com/blog/how-distributional-balances-security-with-insights</link><guid isPermaLink="true">https://distributional.com/blog/how-distributional-balances-security-with-insights</guid><description>Security architecture for AI testing: in-environment deployment, no call-home, namespace access control, and secure integration. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Thu, 26 Jun 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Distributional is designed with enterprise security needs in mind. Our goal has always been to create a testing platform that is secure and compliant—one that protects user privacy and adheres to regulatory requirements without compromise—while still enabling teams to gain insights from their data.&lt;/p&gt;
&lt;p&gt;To achieve this, we’ve architected our product to be private and secure by design. There are several aspects of our product that make it possible for us to work with the most security conscious oriented teams in the world—here are a few of the key elements.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/how-distributional-balances-security-with-insights/685b3613a7ffa1a4814cb978_AD_4nXd7GSQ-5HUaoSZLkTtc5NJTe5-lcC0BT9aPIFWfMPwICPNI7oKfHoygmaaBD4SQ7q8bCCzWZ-dCyCACy7kc3zsTd2iibgVGzwIngkJxm9PiGPS5AmZXn7G--UEm-zZEB4m7waX98w.jpeg&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Distributional is designed to seamlessly integrate within customer infrastructure, securely coexisting with data.&lt;/p&gt;
&lt;h2 id=&quot;deploying-in-a-secure-environment&quot;&gt;&lt;a href=&quot;#deploying-in-a-secure-environment&quot;&gt;Deploying in a secure environment&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional’s platform can be seamlessly deployed within a customer’s environment—whether on-premises or within a VPC—ensuring they maintain full control over data access and security. For teams that prefer even greater autonomy, Distributional also offers a lightweight version of the platform that can be run locally.&lt;/p&gt;
&lt;h2 id=&quot;no-call-home-functionality&quot;&gt;&lt;a href=&quot;#no-call-home-functionality&quot;&gt;No call-home functionality&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;To further ensure that sensitive information remains completely under a customer’s control,&lt;/p&gt;
&lt;p&gt;Distributional is designed with a strict no call-home policy. Once installed in a customer’s environment, it operates entirely within their secure infrastructure, with no data being sent or received outside of their system, aside from automated fully-configurable alerting.&lt;/p&gt;
&lt;h2 id=&quot;secure-access-control&quot;&gt;&lt;a href=&quot;#secure-access-control&quot;&gt;Secure access control&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional is deployed with a namespace architecture, which provides secure and flexible access management. Namespaces serve as isolated partitions within the platform, allowing administrators to control who can access the specific Distributional projects aligned with individual AI apps and any relevant data.&lt;/p&gt;
&lt;h2 id=&quot;integration-with-existing-models&quot;&gt;&lt;a href=&quot;#integration-with-existing-models&quot;&gt;Integration with existing models&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional uses LLM-as-Judge to evaluate whether AI outputs meet the intended task. Teams can connect their own securely hosted model or use Distributional’s default. This approach ensures that even judgment and evaluation maintain the same high standard of privacy and security as core applications. Data never leaves the secure environment, and all models operate under existing security protocols.&lt;/p&gt;
&lt;h2 id=&quot;integration-with-existing-data&quot;&gt;&lt;a href=&quot;#integration-with-existing-data&quot;&gt;Integration with existing data&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Since most customers already maintain their AI production logs in a centralized, secure location, Distributional’s platform is designed to integrate directly with these data stores and automatically fetch the AI production logs. Only the data explicitly approved will be accessed or processed by Distributional. Customers retain full control over which logs leave their storage environment, ensuring that sensitive information remains secure and compliant with their data protection policies.&lt;/p&gt;
&lt;h2 id=&quot;secure-standard-apis&quot;&gt;&lt;a href=&quot;#secure-standard-apis&quot;&gt;Secure, standard APIs&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional offers a robust set of standardized APIs designed for secure, reliable, and scalable integration. All data exchanged with the platform is encrypted using HTTPS, and the APIs are optimized for high-throughput, low-latency operations suitable for production environments.&lt;/p&gt;
&lt;h2 id=&quot;secure-and-insightful-testing-through-metrics&quot;&gt;&lt;a href=&quot;#secure-and-insightful-testing-through-metrics&quot;&gt;Secure and insightful testing through metrics&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional’s platform takes a unique approach to testing by extracting measurable metrics derived from text data using advanced analysis, rather than focusing on just the raw text inputs and outputs. This method offers a critical advantage—ensuring these metrics cannot be used to reconstruct the original text. This allows any sensitive information in the raw text to remain protected. Customers have the flexibility to only rely on these derived metrics for testing, rather than uploading text data. This ensures teams are able to still garner actionable insights, without the risk of exposing sensitive user information.&lt;/p&gt;
&lt;h2 id=&quot;download-the-full-tech-paper&quot;&gt;&lt;a href=&quot;#download-the-full-tech-paper&quot;&gt;Download the full tech paper&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Distributional’s platform is designed to give customers complete control and optionality over their data and integrations. The platform can adapt to specific security, privacy, and regulatory needs, while still providing the quality and insights necessary to test AI applications.&lt;/p&gt;
&lt;p&gt;Learn more about Distributional’s security features by &lt;a href=&quot;http://www.distributional.com/papers/testing-ai-without-risk&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;downloading the full tech paper&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Model Evaluation: From ML to GenAI</title><link>https://distributional.com/blog/model-evaluation-from-ml-to-genai</link><guid isPermaLink="true">https://distributional.com/blog/model-evaluation-from-ml-to-genai</guid><description>Why GenAI evaluation differs from classical ML: multi-component coupling, non-stationarity, non-determinism, and where distributional testing fits. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Thu, 05 Jun 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;As AI systems grow more complex, the methodologies used to evaluate them must evolve accordingly. Traditional machine models are largely deterministic—the same input yields the same output. Machine learning has a long history of model evaluation and benchmarking, making it relatively straightforward to evaluate performance of production ML systems over time. On the other hand, assessing the performance of generative AI systems introduces new complexities.&lt;/p&gt;
&lt;p&gt;Modern AI applications, which may contain GenAI models such as GPT-4o or Claude 4 Sonnnet, have a number of complex properties, such as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI applications are &lt;strong&gt;multi-component systems&lt;/strong&gt; where changes in one part can affect others in unexpected ways. For instance, a change in the vector database could affect the LLM’s responses, or updates to a feature pipeline could impact the machine learning model’s predictions.&lt;/li&gt;
&lt;li&gt;AI applications are &lt;strong&gt;non-stationary&lt;/strong&gt;, meaning their behavior changes over time even if the code doesn’t change. This happens because the world they interact with changes—new data comes in, language patterns evolve, and third-party models get updated. A test that passes today might fail tomorrow, not because of a bug, but because the underlying conditions have shifted.&lt;/li&gt;
&lt;li&gt;AI applications are &lt;strong&gt;non-deterministic&lt;/strong&gt;. Even with the exact same input, they might produce different outputs each time. Think of asking an LLM the same question twice—you might get two different, but equally valid, responses. This makes it impossible to write traditional tests that expect exact matches.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;evaluating-genai-systems&quot;&gt;&lt;a href=&quot;#evaluating-genai-systems&quot;&gt;Evaluating GenAI systems&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The evaluation metrics for GenAI systems look very different than in ML. They are functions that take text as input, such as BLEU or ROUGE, though they seek to answer many of the same questions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Accuracy: Does it give the correct answer?&lt;/li&gt;
&lt;li&gt;Factuality: Are its claims true?&lt;/li&gt;
&lt;li&gt;Helpfulness: Is it useful to the user?&lt;/li&gt;
&lt;li&gt;Safety: Does it avoid harmful or biased outputs?&lt;/li&gt;
&lt;li&gt;Generalization: Can it handle tasks it wasn&amp;#39;t explicitly trained on?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Another way evaluation of GenAI systems differs is that humans (and their opinions) are often involved through the creation of golden—or validation—datasets. These are limited because golden datasets evaluated in development might not reflect the true distributions of the input data (how users interact with the system) over time, and may only reflect certain conditions that were anticipated, whereas real life usage may have many wild deviations from that.&lt;/p&gt;
&lt;h2 id=&quot;distributional-testing&quot;&gt;&lt;a href=&quot;#distributional-testing&quot;&gt;Distributional testing&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;GenAI systems need to be evaluated through distributional testing. In contrast to ML, where tracking the distribution of a single predicted value is very straightforward, the output of LLMs is text-based (structured outputs can contain other metadata) and there are many derived metrics from LLM input/output which can be tracked over time. Some of these metrics include word count, toxicity, ROUGE and BLEU scores (check out a list of standard off-the-shelf LLM eval metrics &lt;a href=&quot;https://docs.dbnl.com/reference/python-sdk/eval-module/dbnl.eval.metrics&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;here&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;).&lt;/p&gt;
&lt;p&gt;Distributional helps teams test input/output distributions over time, gaining insight on exactly how, where, and why their GenAI systems have shifted. &lt;a href=&quot;https://www.distributional.com/get-started&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Get in touch with our team&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; to learn more about Distributional.&lt;/p&gt;
&lt;p&gt;To learn more about model evaluation, watch the full live talk recording below.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=S54cyxg4_q4&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Designing for speed and scale: Integrating Distributional within an existing environment</title><link>https://distributional.com/blog/designing-for-speed-and-scale-integrating-distributional-within-an-existing-environment</link><guid isPermaLink="true">https://distributional.com/blog/designing-for-speed-and-scale-integrating-distributional-within-an-existing-environment</guid><description>How the platform integrates within existing environments: cloud and VPC deployment, storage integration, and data-format agnosticism. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 04 Jun 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Distributional’s adaptive testing platform is designed to support the scale and efficient processing necessary for AI teams to continuously define, understand, and improve AI application behavior. By implementing continuous adaptive testing, teams are able to quantifiably detect and understand when there are deviations from the desired behavior of their AI applications. This in turn helps enterprises bridge the &lt;a href=&quot;https://distributional.com/blog/bridging-the-ai-confidence-gap-with-adaptive-behavioral-testing&quot;&gt;AI Confidence Gap&lt;/a&gt;, so they can productize higher value applications, confidently keep them in production, and achieve the gains that GenAI promises.&lt;/p&gt;
&lt;p&gt;To help teams quickstart today and scale tomorrow, we designed Distributional’s platform to easily integrate within a customer’s environment. In a previous blog, we introduced &lt;a href=&quot;https://distributional.com/blog/understanding-distributionals-platform-architecture&quot;&gt;Distributional’s platform architecture&lt;/a&gt;. In this article, we’ll cover how Distributional integrates within a customer’s existing environment.&lt;/p&gt;
&lt;h2 id=&quot;integrating-within-a-customers-environment&quot;&gt;&lt;a href=&quot;#integrating-within-a-customers-environment&quot;&gt;Integrating within a customer’s environment&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The Distributional Platform is designed to sit within a customer’s existing data infrastructure, rather than be a separate silo to manage and sync. It can easily be deployed in the existing cloud environment and integrate with the existing storage system. This makes it easy to fit seamlessly within any AI platform. By leveraging industry-standard components, it also ensures that customers can take advantage of the managed cloud service versions of each for even easier management.&lt;/p&gt;
&lt;p&gt;In addition to being fairly agnostic with regard to where the data lives, the platform is also agnostic on what the data looks like. Any existing logs, traces, or other evaluation metrics can be used as inputs. The more data provided, the more context the platform has to develop a comprehensive understanding of behavior for testing.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/designing-for-speed-and-scale-integrating-distributional-within-an-existing-environment/68237f0ca7e35c89cb8dec29_AD_4nXd4iZ9Wp_-MG3WC5yMY51PHMhP2CVP8VZc432WrpKgg1YHHqUzHPE4vhEbBw2rjqxjvXasXy2X_xPODHIam-SEuuq2zQHNJ5F_r6lOujvmm4YsUF0deAniUKG09nSC_eOvv5A_yRA.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Customers can then use their preferred orchestration tool to define the ingest schedule for how often new data is sent to Distributional. For example, a customer could schedule daily updates through Airflow to ensure their models haven&amp;#39;t drifted by feeding a fixed set of inputs into their AI app to generate a dataset of examples to be uploaded to Distributional as a run and tested for change.&lt;/p&gt;
&lt;p&gt;For authentication, the Distributional platform uses OpenID Connect (OIDC), again ensuring seamless integration with any preferred identity providers. In addition to authentication, the platform also has a native permission model using role-based access controls for further security.&lt;/p&gt;
&lt;p&gt;Overall, the SDK is designed for broad extensibility. It’s built in Python, making it flexible for customers to integrate with other preferred tools or existing processes. Distributional is also continuing to expand the native integrations built into the platform to make this even more seamless for customers. This allows customers to seamlessly integrate Distributional with their AI platforms, while still allowing for portability and adaptability as these platforms continue to mature.&lt;/p&gt;
&lt;p&gt;Ultimately, this results in a platform that is easy to deploy and manage, integrates within an existing environment and preferred tooling, and is built for enterprise scale and secure usage.&lt;/p&gt;
&lt;h2 id=&quot;access-the-full-tech-paper&quot;&gt;&lt;a href=&quot;#access-the-full-tech-paper&quot;&gt;Access the full tech paper&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;If you’re interested in learning more about Distributional’s platform and architecture, check out the &lt;a href=&quot;https://www.distributional.com/papers/introduction-to-distributionals-platform&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;full tech paper&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;. If you’re interested in trying out Distributional’s adaptive testing platform, &lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;reach out to the team&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; and we’d be happy to get you set up.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Understanding Distributional’s Platform Architecture</title><link>https://distributional.com/blog/understanding-distributionals-platform-architecture</link><guid isPermaLink="true">https://distributional.com/blog/understanding-distributionals-platform-architecture</guid><description>Distributional&apos;s platform architecture: integrate with existing storage, scheduled batch ingestion, and analytic-style processing of AI logs. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Thu, 22 May 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Adaptive testing is critical to support production-grade AI applications at scale. This enables AI teams to leverage behavioral distributions to create a comprehensive definition of an app’s desired behavior that can be refined over time. Thus enabling teams to quantifiably detect and understand when there are deviations from that desired behavior.&lt;/p&gt;
&lt;p&gt;Using Distributional’s platform, customers are able to productize higher value applications, and keep them in production, all while minimizing risk to the business with the confidence that these applications are behaving and will continue to behave as desired.&lt;/p&gt;
&lt;p&gt;Distributional provides an enterprise platform for adaptive testing. To support the scale and performance needs for AI teams to continuously test, understand, and improve application behavior, the platform needed to handle large, context-rich datasets as well as perform fast, analytic-style processing of that data. Let’s go into more detail about how the platform is architected to support this.&lt;/p&gt;
&lt;h2 id=&quot;platform-architecture&quot;&gt;&lt;a href=&quot;#platform-architecture&quot;&gt;Platform architecture&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Importantly, Distributional’s platform does not try to replace customers&amp;#39; existing data storage systems. It is purposefully designed to integrate with existing data infrastructure, where all the raw logs and LLM traces for the AI applications already are stored. Based on the customer-defined ingestion schedule, Distributional is then able to efficiently process this data in batch, with inherently no limits on scale. This enables the platform to run much more complex analytics against the data, while maintaining its rich context, to derive a comprehensive set of metrics and statistics about the applications, as well as calculate changes to these over time.&lt;/p&gt;
&lt;p&gt;Additionally, unlike systems designed for pure logging and monitoring, users can update or change the data provided to Distributional. This gives them the ability to do things like adding in missing data points and retrying the processing job, expanding the type of data and context provided, or even recalculating metrics based on data from a past point in time.&lt;/p&gt;
&lt;p&gt;Since access to this data is critical, customers deploy and manage the Distributional platform in their private cloud environment. This prevents Distributional from becoming yet another technology silo within a customers overall architecture and allows the platform to live where the data already resides, preventing duplicate systems of record or concerns about moving data outside of the existing secure environment. To support this though, we purposefully architected it with as few dependencies as possible and leveraged industry-standard systems to make it as easy to deploy and manage as possible.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/understanding-distributionals-platform-architecture/681163b9c370ea3e44d4a070_AD_4nXdHE2nw3RW_dWefCyjbKiRAR2YWNTDKHwOF5jfjyuQm5SInMFwOsuCjjWeFvTgHgqBUrqbAHmtgr3LY87oqed53FzgxgkvRdPn8qI11WzzqbMtsYX-PcDwNyhlqss7KWG7mcsz94Q.png&quot; alt=&quot;Distributional&apos;s platform architecture&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Distributional&amp;#39;s platform architecture.&lt;/p&gt;
&lt;p&gt;At a high level, the platform consists of an API service, a UI service, a worker service, a messaging queue, and a database. All three services run on Kubernetes with Redis and PostgreSQL being used for the messaging queue and database, respectively. The data is stored in an object store with support for AWS S3, Google Cloud Storage (GCS), and Azure Blob Storage. This results in only four external dependencies that all leverage industry-standards our customers are already familiar with.&lt;/p&gt;
&lt;p&gt;The core of the platform is the API service, with users able to interact with it via the user interface or SDK for more programmatic access. When the API service is tasked with a workload such as computing metrics on raw data or evaluating different tests, the messaging queue relays the work to the worker service for compute processing and can scale out these resources or adjust the type of compute as needed depending on the size of job or concurrent requests. For the users, this design abstracts away the complexity of managing storage and compute resources, so they can focus on getting the answers they need from their application’s data.&lt;/p&gt;
&lt;p&gt;If you’re interested in learning more about Distributional’s platform and architecture, check out the &lt;a href=&quot;https://www.distributional.com/papers/introduction-to-distributionals-platform&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;full paper&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;. If you’re interested in trying out Distributional’s adaptive testing platform, &lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;reach out to the team&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; and we’d be happy to get you set up.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Adaptive testing for AI confidence</title><link>https://distributional.com/blog/adaptive-testing-for-ai-confidence</link><guid isPermaLink="true">https://distributional.com/blog/adaptive-testing-for-ai-confidence</guid><description>A three-step methodology for adaptive AI testing: define desired behavior, understand change, continuously adapt as the state evolves. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 21 May 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;While &lt;a href=&quot;https://distributional.com/blog/bridging-the-ai-confidence-gap-with-adaptive-behavioral-testing&quot;&gt;confidence may be hindering AI progress&lt;/a&gt; at your organization, there is hope. The answer lies in rethinking testing. You can no longer rely solely on the limited and static performance-based tests that have worked for teams in the past. Instead, you need to move to adaptive testing to support production-grade AI applications at scale.&lt;/p&gt;
&lt;p&gt;Adaptive testing leverages behavioral distributions to create a comprehensive definition of your app’s desired behavior that you refine over time. This enables you to quantifiably test and investigate when there are deviations from that desired behavior. Ultimately you are able to define a desired application state with these behavioral tests, understand as this state evolves, and be able to adapt the tests to reflect this changing state.&lt;/p&gt;
&lt;p&gt;But what does it take to get there? Let’s take a closer look at the three steps necessary to bring adaptive testing to your AI apps.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/adaptive-testing-for-ai-confidence/67fe80d6da9e11fd73a728fd_Adaptive-Testing-Workflow_v1.avif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;The three steps (and continuous adaptive process) of bringing adaptive testing to your AI apps.&lt;/p&gt;
&lt;h3 id=&quot;step-one-define-desired-behavior&quot;&gt;&lt;a href=&quot;#step-one-define-desired-behavior&quot;&gt;&lt;strong&gt;Step one: Define desired behavior&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;First, you need to quantify, define, and test the behavior of your AI apps, so that you can continuously understand how these apps behave in production both today and in the future.&lt;/p&gt;
&lt;p&gt;Comprehensively quantifying behavior requires taking into account a richer, more complete set of your application’s attributes to test against. The more attributes, the better the understanding of desired behavior. And the better that understanding, the more likely it is you’ll know when there’s a change to that behavior. Given the breadth of coverage necessary, having an adaptive testing solution to help you automate this is key. It should be able to leverage existing log data or golden datasets for your app to automatically generate a robust set of statistical distributions of these attributes, resulting in a complete, unique fingerprint of your app’s behavior. This allows you to easily define what behaviors or behavioral changes you want to be alerted to over time and explicitly test for them. And then adapt your behavioral tests to best align with your desired behavior over time.&lt;/p&gt;
&lt;h3 id=&quot;step-two-understand-changes-in-behavior&quot;&gt;&lt;a href=&quot;#step-two-understand-changes-in-behavior&quot;&gt;&lt;strong&gt;Step two: Understand changes in behavior&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;With AI, change is inevitable. So once you’re able to test whether your app is behaving within your definition of desired behavior, next you need to understand the changes that led to this. What actually changed? What was the impact on behavior? And what caused the change?&lt;/p&gt;
&lt;p&gt;With adaptive testing, you get the depth necessary to quickly investigate any changes to behavior and decide what action to take. When you see there’s been a change to your app’s behavior, you’re able to dig into the exact attributes that changed and pinpoint the specific results that most contributed to this change. Adaptive testing maintains this lineage of insights at every layer of your app, even with increasing complexity or scale, to give you a deeper understanding of what’s changing and its cascading impact. Armed with that, you’re able to shift from passive observation to insights you can action upon. You’re able to fully assess whether this change is acceptable or whether there’s an issue that needs to be resolved. For instance, if the behavioral change is due to a shift in production usage, you may choose to adjust your tests to better represent this new state of desired behavior. But if it is related to an issue due to model drift, you’re able to share the results of your root cause analysis with your development teams and work with them to resolve the issue with minimal impact to production.&lt;/p&gt;
&lt;h3 id=&quot;step-three-continuously-improve-with-these-changes-in-behavior&quot;&gt;&lt;a href=&quot;#step-three-continuously-improve-with-these-changes-in-behavior&quot;&gt;&lt;strong&gt;Step three: Continuously improve with these changes in behavior&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Finally, you’re able to continuously improve your app and overall development lifecycle with these changes in behavior. For any newly discovered change to behavior, you’ll either update your tests to encompass this behavior if it is acceptable or update your app to address the behavior if it is undesired. Then you’ll continue with testing your app’s behavior—either running the new set of tests that now represent desired state, or testing whether your updates to the app adhere to desired state. Thus, adaptive testing becomes the way you can get a snapshot of behavioral state for an AI app and continuously confirm it is what is desired over time.&lt;/p&gt;
&lt;p&gt;By automating this process, you free up time to develop and push new updates to your apps, all with the confidence that it won’t impact production behavior. You can test and assess the impact of swapping to new models or adding new components and functionality, while leveraging the same suite of adaptive tests and definition of behavior. By having this persistent and shared definition of behavior, your team can also collaborate more efficiently across the development lifecycle to ensure the resulting apps can keep pace with the changing market needs, and spend time creating the AI apps that will truly move the needle for your business.&lt;/p&gt;
&lt;h3 id=&quot;adaptive-testing-is-the-missing-piece-for-ai&quot;&gt;&lt;a href=&quot;#adaptive-testing-is-the-missing-piece-for-ai&quot;&gt;&lt;strong&gt;Adaptive testing is the missing piece for AI&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Adaptive testing can help your team mature both with how you develop AI apps but also with the types of applications that make it into production—and stay in production.&lt;/p&gt;
&lt;p&gt;But as with any technology solution, it cannot be yet another silo. This is a critical backbone for your overall AI technology stack. To get the full benefits, an adaptive testing solution needs to integrate into your existing processes and technologies. Such as integrating with the platforms that already house your application logs and existing eval metrics, or hooking into the observability and alerting tools already in place. These will help you get started faster and minimize disruption to your team.&lt;/p&gt;
&lt;p&gt;The right solution must meet your team where you are today and act as a catalyst to help you safely and confidently scale development and usage. To learn more about Distributional’s adaptive testing solution and see if it’s the right fit for you, reach out to us &lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;here&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Distributional Simplifies Adaptive Testing with Similarity Index and Key Insights</title><link>https://distributional.com/blog/distributional-simplifies-adaptive-testing-with-similarity-index-and-key-insights</link><guid isPermaLink="true">https://distributional.com/blog/distributional-simplifies-adaptive-testing-with-similarity-index-and-key-insights</guid><description>Why static golden-dataset testing breaks, and how Similarity Index and Key Insights pinpoint and explain behavioral change. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 21 May 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;While seemingly everyone is experimenting with AI, only a few have been able to bridge the &lt;a href=&quot;https://distributional.com/blog/bridging-the-ai-confidence-gap-with-adaptive-behavioral-testing&quot;&gt;AI confidence gap&lt;/a&gt; into production and handling real world usage. And even fewer are confident enough to tackle the most valuable – and risky – AI use cases. We believe &lt;a href=&quot;https://distributional.com/blog/adaptive-testing-for-ai-confidence&quot;&gt;adaptive testing&lt;/a&gt; is the key. But why is it that current approaches to testing are breaking down?&lt;/p&gt;
&lt;p&gt;It’s not that teams aren’t doing testing. In fact, nearly all enterprise teams that we’ve met with have implemented some form of testing, especially during the development of their AI applications. But this testing tends to be incomplete, static, and hard to act upon. For example, teams may start with testing different prompts and comparing individual responses one-by-one. The best performing responses in turn get annotated and become a golden dataset. And this golden dataset is used to test different models and select the best performing one.&lt;/p&gt;
&lt;p&gt;Once in production, teams may also start tracking some aggregate summary statistics, as well as user feedback on the quality of responses. These tests are run periodically on a subset of responses just to be sure. Until the response performance degrades. And these teams are forced back to square one just to figure out what’s wrong, resulting in entirely new research and development cycles to try and get performance back on track. For AI applications, testing needs to go beyond simply understanding performance and provide an understanding of behavior as a whole.&lt;/p&gt;
&lt;p&gt;“&lt;em&gt;GenAI has created unique challenges that aren’t well handled by existing MLOps platforms and workflows. Most notably, enterprises seek ways to ensure that GenAI systems behave as expected and don’t introduce unpredictable behavior that could result in reputational harm, poor user experiences, or costly business consequences,&lt;/em&gt;” says Sam Charrington, Creator and Host of the TWIML AI Podcast. &lt;em&gt;“New approaches to system profiling and testing are required to meet this need, and Distributional’s adaptive testing solution is purpose-built to solve this for enterprise teams.&lt;/em&gt;”&lt;/p&gt;
&lt;p&gt;Helping these teams understand the full behavior of their AI apps so they can deploy with confidence is why we created Distributional, the first enterprise platform for adaptive testing. Using Distributional’s platform, customers are able to maximize the uptime of AI apps, resolve issues faster, and unlock the development of higher value applications. All while minimizing risk to the business with the confidence that these applications are behaving - and will continue to behave - as desired.&lt;/p&gt;
&lt;p&gt;Today, we’re excited to introduce some new capabilities to help enterprise teams more easily implement adaptive testing. But let’s first take a deeper look at the adaptive testing workflow to better understand where these features will fit in.&lt;/p&gt;
&lt;h3 id=&quot;adaptive-testing-workflow&quot;&gt;&lt;a href=&quot;#adaptive-testing-workflow&quot;&gt;Adaptive testing workflow&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/distributional-simplifies-adaptive-testing-with-similarity-index-and-key-insights/67fee2dcb3a8668bf6e5da1b_workflow.avif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Distributional&amp;#39;s adaptive testing workflow&lt;/p&gt;
&lt;p&gt;Distributional’s platform is focused on providing a simple and automated workflow to help AI teams across enterprises continuously define, understand, and improve AI application behavior.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Define&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Distributional first helps teams define the desired behavior for their applications. The platform automatically creates a behavioral fingerprint using the app’s runtime logs as well as any existing development metrics. Distributional generates associated tests to be able to detect changes in that behavior over time. Teams can also use this behavioral fingerprint to specify behaviors they do or do not want the application to exhibit at a statistical level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Understand&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Once that behavior is defined, teams are able to use the platform to understand changes in behavior and deviations from desired behavior as these apps are being used in production. They get alerted when there are changes to app behavior, understand what is changing, and pinpoint at any level of depth what is causing the change to quickly take appropriate action.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Improve&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Finally, these teams continuously improve both their tests and their app based on any changes they observe. By easily adding, removing, or recalibrating tests over time, teams now have a dynamic and accurate representation of desired state to test new models, roll out new upgrades, or accelerate new app development.&lt;/p&gt;
&lt;p&gt;This workflow is game changing in driving more consistent and predictable app behavior. And we will constantly innovate in the product to ensure teams have what they need to be confident in their AI, especially as they support and scale more apps. Today’s innovations continue that commitment, providing clarity into what changed and whether an action is needed.&lt;/p&gt;
&lt;h3 id=&quot;introducing-the-similarity-index--key-insights&quot;&gt;&lt;a href=&quot;#introducing-the-similarity-index--key-insights&quot;&gt;Introducing the Similarity Index &amp;amp; Key Insights&lt;/a&gt;&lt;/h3&gt;
&lt;h4 id=&quot;pinpointing-change-with-similarity-index&quot;&gt;&lt;a href=&quot;#pinpointing-change-with-similarity-index&quot;&gt;Pinpointing change with Similarity Index&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;While it’s necessary to take into account a much more robust set of attributes to get a comprehensive, testable understanding of behavior, teams need to balance this with the ability to quickly understand “is the current behavior of my AI app similar to what I want it to be? If not, where is it least similar and do I care?”. As these teams scale to more usage and app complexity, the ability to easily answer these questions is even more critical. Which is why we developed the Similarity Index and Key Insights.&lt;/p&gt;
&lt;p&gt;The Similarity Index (Sim Index) is a single numerical value — between 0 and 100 — that quantifies how much an application or subsets of an application has changed between two points in time. Think of it as a signal that gives you an instant read on whether your app is behaving consistently, or if something has meaningfully shifted.&lt;/p&gt;
&lt;p&gt;Sim Index operates across three levels:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Application-level&lt;/strong&gt;: how much your app as a whole has drifted&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Column-level&lt;/strong&gt;: which specific inputs, outputs, or intermediate data have changed&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Metric-level&lt;/strong&gt;: what specific properties or evals of the column level data (e.g. readability, accuracy) are driving the change&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The Sim Index is automatically computed on every test session with no setup required and can be integrated with existing alerting tools like Pagerduty so you can get notified when there is a change in the value.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/distributional-simplifies-adaptive-testing-with-similarity-index-and-key-insights/67feec0051a787e40d347a35_Sim_Index_v4.gif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Similarity Index&lt;/p&gt;
&lt;p&gt;When there is a drop in the Sim Index — say from 94 to 46 — you immediately know that there has been a significant change. To help you quickly understand what has changed is where Key Insights come in.&lt;/p&gt;
&lt;h4 id=&quot;understanding-change-with-key-insights&quot;&gt;&lt;a href=&quot;#understanding-change-with-key-insights&quot;&gt;Understanding change with Key Insights&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Alongside every Sim Index are Key Insights, which are designed to give you an actionable interpretation of what has changed at a glance. The Distributional platform automatically generates these human-readable summaries that tell you exactly what changed and why it matters.&lt;/p&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;p&gt;&lt;em&gt;“Answer similarity has dropped significantly — response length distribution substantially drifted to the right and readability decreased by 20%.”&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Each Insight is tied to the relevant metric-level details and includes visual comparisons and historical context so you get instant clarity and can take immediate action, such as triaging issues or validating hypotheses. Even more powerful is the ability to create new tests and set thresholds on critical metrics with a single click, so you can continuously add behavioral test coverage that best represents your desired behavior for your specific application.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/distributional-simplifies-adaptive-testing-with-similarity-index-and-key-insights/67fef1a8c7eb0738b6be5950_Key_Insights_v3.gif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Key Insights&lt;/p&gt;
&lt;p&gt;Together, Similarity Index and Key Insights give teams:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;broad view&lt;/strong&gt; of application behavior and how it’s changing over time&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;fast, guided path to root cause&lt;/strong&gt;, with the ability to drill down from the application-level, down to the column- and even metric-level.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These features are now available in Distributional’s platform and are automatically included as part of the default setup, making it easier than ever to gain confidence in the behavioral stability of your AI apps from day one. To see a full demo of these features in action, check out this &lt;a href=&quot;https://youtu.be/F2e9_UxyrVs&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;video&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;If you’re interested in trying out these new capabilities for yourself, &lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;reach out to the team&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; and we’d be happy to get you set up.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Driving confidence in AI: Understanding deployment testing through model version updates</title><link>https://distributional.com/blog/driving-confidence-in-ai-understanding-deployment-testing-through-model-version-updates</link><guid isPermaLink="true">https://distributional.com/blog/driving-confidence-in-ai-understanding-deployment-testing-through-model-version-updates</guid><description>Deployment testing through LLM version updates: metric choices, vibe checks versus statistical perspective, and automating the comparison. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 16 Apr 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Building on my previous &lt;a href=&quot;https://distributional.com/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example&quot;&gt;blog post&lt;/a&gt;, I’d like to share more examples to help readers develop intuition around AI testing. This time, rather than focusing on RAG apps, we’ll dive into deployment testing. The goal is to illustrate how deployment testing can provide deeper insights into the consistency of an AI-powered application after it reaches production.&lt;/p&gt;
&lt;p&gt;At its core, deployment testing is about understanding the consistency of an AI application over time. This can range from a narrow focus of continuously testing a single LLM-endpoint, to the broader scope of continuously testing an entire AI application.&lt;/p&gt;
&lt;p&gt;Regardless of the scope, the goal is to extract a meaningful signal that indicates whether or not  something is changing. Interestingly, testing the full AI application is often easier, because we generally have a clearer understanding of what the application is supposed to do. In contrast, testing a single LLM-endpoint can be more challenging, as its expected behavior is harder to define.&lt;/p&gt;
&lt;p&gt;There are multiple ways to extract meaningful signals. I personally like to extract signals through summarization. A simple system diagram for a summarization application is shown below:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[{&amp;quot;text&amp;quot;: &amp;quot;Inflation has been rising in many countries, driven by supply chain disruptions, increased demand, and higher energy prices. Central banks are grappling with how to respond, balancing the need to control inflation without stifling economic recovery. Experts warn that high inflation could lead to reduced purchasing power and economic instability.&amp;quot;,&amp;quot;reference_summary&amp;quot;: &amp;quot;Rising inflation, driven by supply chain issues and higher energy prices, challenges central banks in balancing control with economic recovery.&amp;quot;}]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And similarly from the Technology category:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[{&amp;quot;text&amp;quot;: &amp;quot;Edge computing is transforming how data is processed by moving computation closer to the data source. This reduces latency and enhances performance for applications like autonomous vehicles and real-time data analysis. However, managing and securing decentralized data remains a significant challenge.&amp;quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; &amp;quot;reference_summary&amp;quot;: &amp;quot;Edge computing reduces latency by processing data closer to the source, enhancing performance, though managing and securing decentralized data is challenging.&amp;quot;}]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Obviously, if we were to test a full AI-powered application, this Golden Dataset should reflect the tasks which we wanted the application to complete. However, for testing a single LLM-endpoint, we instead only need to come up with a meaningful Golden Dataset that ideally reflects the business relying on the LLM-endpoint.&lt;/p&gt;
&lt;h3 id=&quot;what-about-the-metrics&quot;&gt;&lt;a href=&quot;#what-about-the-metrics&quot;&gt;What about the metrics?&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;When evaluating behavior, relying on a single metric is never sufficient. However, since the focus of this blog is on building intuition, we’ll keep it simple and examine just two key metrics.&lt;/p&gt;
&lt;p&gt;For this example, we’ll look at the BLEU and the ROUGE score, which are both well-established NLP-metrics that show differences between two pieces of text. One could imagine expanding these two with all sorts of metrics like toxicity-level and readability scores.&lt;/p&gt;
&lt;h3 id=&quot;and-what-about-the-llm-endpoint&quot;&gt;&lt;a href=&quot;#and-what-about-the-llm-endpoint&quot;&gt;And what about the LLM-endpoint?&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;From countless conversations I’ve had, there are three primary ways teams access LLMs:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Third-Party Endpoints&lt;/strong&gt; – Services like ChatGPT and Together AI that provide external API access.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Self-Hosted&lt;/strong&gt; – Teams running their own LLMs on dedicated infrastructure.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Company-Hosted&lt;/strong&gt; – Organization-wide initiatives where LLMs are deployed on internal infrastructure for company-wide use.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Each approach comes with its own trade-offs. For this blog, we’ll focus on the company-hosted model and act as an engineering team using a Llama LLM-endpoint provided through a company-wide GenAI platform. This means that we have very little control over the LLM-endpoint, and that we’ll be using what is being provided to us from a different team&lt;/p&gt;
&lt;p&gt;Additionally, we’ll be looking at an 11-month time period from May 2024 to March 2025.&lt;/p&gt;
&lt;h3 id=&quot;do-vibe-checks-work&quot;&gt;&lt;a href=&quot;#do-vibe-checks-work&quot;&gt;Do vibe checks work?&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;We’ve found that most companies start with basic vibe checks. So, let’s begin by examining a single piece of text alongside its reference summary.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[{&amp;quot;text&amp;quot;: &amp;quot;Edge computing is transforming how data is processed by moving computation closer to the data source. This reduces latency and enhances performance for applications like autonomous vehicles and real-time data analysis. However, managing and securing decentralized data remains a significant challenge.&amp;quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; &amp;quot;reference_summary&amp;quot;: &amp;quot;Edge computing reduces latency by processing data closer to the source, enhancing performance, though managing and securing decentralized data is challenging.&amp;quot;}]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Next, we can query our LLM-endpoint to retrieve and log the summary generated for the above example each month.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;May 2024 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing reduces latency and enhances performance by processing data closer to the source, ideal for applications like autonomous vehicles and real-time analysis.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;June 2024 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing reduces latency and enhances performance by moving computation closer to data sources, ideal for applications like autonomous vehicles and real-time analysis. However, managing and securing decentralized data remains a significant challenge.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;July 2024 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing reduces latency and enhances performance by moving computation closer to data sources, ideal for applications like autonomous vehicles and real-time analysis. However, managing and securing decentralized data remains a significant challenge.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;August 2024 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing transforms data processing by moving computation closer to the data source, reducing latency and enhancing performance for applications like autonomous vehicles and real-time data analysis.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;September 2024 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing is transforming data processing by moving computation closer to the data source, reducing latency and enhancing performance for applications like autonomous vehicles and real-time data analysis.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;October 2024 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing moves computation closer to data sources, reducing latency and enhancing performance for applications like autonomous vehicles and real-time data analysis, while managing and securing decentralized data remains a significant challenge.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;November 2024 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing moves computation closer to data sources, reducing latency and enhancing performance for applications like autonomous vehicles and real-time data analysis, but managing and securing decentralized data remains a significant challenge.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;December 2024 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing moves computation closer to data sources, reducing latency and enhancing performance for applications like autonomous vehicles and real-time data analysis.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;January 2025 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing reduces latency and enhances performance by processing data closer to its source, but managing and securing decentralized data remains a challenge.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;February 2025 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing reduces latency and enhances performance by processing data closer to its source, but managing and securing decentralized data remains a challenge.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;March 2025 – Summary of above example&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;quot;Edge computing reduces latency and enhances performance by processing data closer to its source, but managing and securing decentralized data remains a challenge.&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;We can observe that the generated summaries vary from month to month, which aligns to the non-deterministic nature of an LLM. However, as humans, we can’t easily determine whether these differences are purely due to non-determinism or if other factors are at play.&lt;/p&gt;
&lt;p&gt;This is where a statistical lens can help!&lt;/p&gt;
&lt;h3 id=&quot;a-statistical-perspective&quot;&gt;&lt;a href=&quot;#a-statistical-perspective&quot;&gt;A statistical perspective?&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;With our Golden Dataset in place, we no longer need to rely on simple month-over-month vibe checks. We can instead examine the distributions of key metrics and track their evolution over time.&lt;/p&gt;
&lt;p&gt;This kind of analysis is at the core of what the Distributional platform enables at scale. However, to build intuition, we’ll continue with the simple two-metric example outlined earlier.&lt;/p&gt;
&lt;p&gt;The workflow is as follows:&lt;/p&gt;
&lt;p&gt;Each month, we will:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Generate a summary for each text-reference-summary pair.&lt;/li&gt;
&lt;li&gt;Compute ROUGE and BLEU scores for each reference-summary and generated-summary pair.&lt;/li&gt;
&lt;li&gt;Calculate the average ROUGE and BLEU scores for each subclass in the Golden Dataset.&lt;/li&gt;
&lt;li&gt;Compute the overall average ROUGE and BLEU scores across all data.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;From here, we can visualize how both ROUGE and BLEU scores change month over month for both the entire Golden Dataset and each of the subclasses.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-understanding-deployment-testing-through-model-version-updates/67ff2eff1fcf6edc8270e35d_image1.webp&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Monthly ROUGE score for each subclass and the global average.&lt;/p&gt;
&lt;p&gt;When analyzing the ROUGE score over time, we observe that our AI application is gradually improving. However, its performance isn’t consistent across all subclasses—some are summarized more effectively than others according to our reference answer. Interestingly, as the BLEU score below shows, an improvement in one metric doesn’t necessarily mean all metrics will improve in parallel.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-understanding-deployment-testing-through-model-version-updates/67ff2f1f2e2cdc611532c1c9_image6.webp&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Monthly BLEU score for each subclass and the global average.&lt;/p&gt;
&lt;p&gt;The BLEU score follows a similar trend to the ROUGE score, but at certain points, we observed a regression that wasn’t reflected in the ROUGE metric. This is the exact reason why we want people to think about change in the context of the distributional fingerprint and not just a single metric. The regression is seen on both August – 2024 and September – 2024 for the BLUE score but not for the ROUGE score.&lt;/p&gt;
&lt;h3 id=&quot;what-is-happening&quot;&gt;&lt;a href=&quot;#what-is-happening&quot;&gt;What is happening?&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Why are we seeing an increase in performance over time? From the plots below, we can see that the team serving up the LLM-endpoint have been following the Llama release schedule, and have been constantly updating the underlying LLM-endpoint to the most recent Llama model.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-understanding-deployment-testing-through-model-version-updates/67ff2f47a866b920727f041e_image4.webp&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Monthly ROUGE score for each subclass and the global average including used LLM.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-understanding-deployment-testing-through-model-version-updates/67ff2f63b3e6f0e7005faa24_image3.webp&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Monthly BLEU score for each subclass and the global average including used LLM.&lt;/p&gt;
&lt;p&gt;Whether or not teams are doing this type of model update out in the wild is not the point—there are plenty of arguments for why one should or shouldn’t do it. What I want to show is that even for simple tasks like summarization, we are capable of getting a really strong indication of whether or not something is changing using the right dataset.&lt;/p&gt;
&lt;p&gt;For situations where you are constantly updating models to the newest version, having a system like the above to extract a continuous signal makes a lot of sense. However, for everyone that relies on a model provider for their LLM-endpoints, this is also a simple way to understand whether or not these providers are making changes to their models, and the impact of those changes.&lt;/p&gt;
&lt;h3 id=&quot;is-there-a-signal&quot;&gt;&lt;a href=&quot;#is-there-a-signal&quot;&gt;Is there a signal?&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;One last thing. It’s fair to ask whether everything we’ve seen so far is just a reflection of the inherent non-determinism of LLMs or if there’s actually a meaningful difference between these models. To dig into that, I ran the data through each LLM 10 times to see if the results truly vary or if we’re just seeing randomness in action.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-understanding-deployment-testing-through-model-version-updates/67ff2f8d48e223bb30bb42c9_image2.webp&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Average ROUGE score including standard deviation for each used LLM.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-understanding-deployment-testing-through-model-version-updates/67ff2fa79c0639b1e7e76fbe_image7.webp&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Average BLEU score including standard deviation for each used LLM.&lt;/p&gt;
&lt;p&gt;From the two plots, it is clear that there &lt;em&gt;is&lt;/em&gt; a difference between the four models. But whether or not that difference will have an impact on an AI powered application is a topic that we’ll cover in a later blogpost.&lt;/p&gt;
&lt;h3 id=&quot;automating-deployment-testing&quot;&gt;&lt;a href=&quot;#automating-deployment-testing&quot;&gt;Automating deployment testing&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Distributional is helping companies define, understand, and improve the reliability of their AI applications. The above scenario represents just a subset of the challenges we’re helping to solve.&lt;/p&gt;
&lt;p&gt;Our goal is to help teams develop confidence in the AI applications they’re taking to production, so they can rely on more than just a vibe check to ensure their applications are working as intended.&lt;/p&gt;
&lt;p&gt;Interested in exploring how Distributional can help your organization? &lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Sign up&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; for access to our product or &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;get in touch&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with our team to learn more.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Bridging the AI Confidence Gap with adaptive behavioral testing</title><link>https://distributional.com/blog/bridging-the-ai-confidence-gap-with-adaptive-behavioral-testing</link><guid isPermaLink="true">https://distributional.com/blog/bridging-the-ai-confidence-gap-with-adaptive-behavioral-testing</guid><description>Why AI confidence erodes between development benchmarks and production, and adaptive behavioral testing over distributions as the fix. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 05 Mar 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;Do you know how your AI app is behaving? What about how that behavior might have changed from yesterday? For most teams, the ability to answer these remains out of reach, especially once an application is launched and being used by real users. While teams might be able to create a partial understanding of behavior for these applications during initial development through performance monitoring or benchmarks, their confidence in this understanding erodes by the time the apps are deployed at scale. And this lack of confidence ends up having a cascading impact on which applications actually end up making it to production.&lt;/p&gt;
&lt;p&gt;To unlock the full potential of AI, you need to be confident in their apps’ behavior from development to deployment to continued usage. But why does this feel so out of reach for most of us today?&lt;/p&gt;
&lt;h3 id=&quot;the-confidence-gap&quot;&gt;&lt;a href=&quot;#the-confidence-gap&quot;&gt;&lt;strong&gt;The confidence gap&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Software teams have been shipping with confidence for decades. These teams define how their applications should perform, and then use standard tests and benchmarks to measure and validate whether the code is working as intended across the development lifecycle. When a test fails, there’s clear processes for triage, root cause analysis, and resolution to get it back online quickly, with minimal production impact.&lt;/p&gt;
&lt;p&gt;Of course there’ve been efficiencies added over the years but, for the most part, this has been enough. The code itself doesn’t change that often, and it’s pretty straightforward to test whether the same outputs are produced from a set of inputs. Even as complexity has increased to handle things like ephemeral cloud services, larger scale distributed systems, and increasing dependencies—testing has kept pace to help teams build and deliver more impactful solutions that align with strategic business objectives and help maintain a competitive edge.&lt;/p&gt;
&lt;p&gt;AI broke the mold. Gone are the days of fixed inputs, defined outputs, and predictable changelogs. By design, AI systems are non-deterministic with the same input able to return a variety of potential outputs. Additionally, there are constant shifts across data, usage, models, prompts, etc. from one day to the next. Not only is it challenging to understand when these shifts happen, but also to trace the impact of these shifts back to specific components to resolve or understand the extent they’ve been propagated throughout these systems, especially when there are multiple models daisy chained together.&lt;/p&gt;
&lt;p&gt;AI requires teams to move beyond this fixed understanding of performance monitoring and testing. To reinject confidence and bridge the gap from development to production for AI, it’s time to move to adaptive behavioral testing.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/bridging-the-ai-confidence-gap-with-adaptive-behavioral-testing/67fe808a4235be9e70da83b8_From-Confidence-to-Production-2048x1152.avif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Adaptive behavioral testing gives AI teams the coverage, clarity, and continuity necessary for productionizing AI apps and keeping them in production.&lt;/p&gt;
&lt;h3 id=&quot;build-confidence-through-adaptive-behavioral-testing&quot;&gt;&lt;a href=&quot;#build-confidence-through-adaptive-behavioral-testing&quot;&gt;&lt;strong&gt;Build confidence through adaptive behavioral testing&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;With adaptive behavioral testing, you can establish a more complete understanding of behavior as a whole, and quantify that behavior into something you can test and measure continuously. But how do you get there?&lt;/p&gt;
&lt;p&gt;First, it’s about removing uncertainty as you build your definition of desired behavior and adapt it for the future. You need to take into account a richer, more robust set of properties for your app—not just the inputs and outputs related to what it produced but all of the intermediates related to how it behaved to achieve those outputs. And rather than test these properties against a single threshold or summary statistic, you need to test based on their distributions or range of acceptability.&lt;/p&gt;
&lt;p&gt;Evaluation metrics and benchmarks can be a great place to start, but alone aren’t enough. Your teams will never be able to create evals for every possible edge case that could happen in production. And even if that were possible, it would only represent what’s true today, not what could happen in the future. Plus, these benchmarks are designed to only measure against the end outputs, but stop short of measuring and understanding the behavior to get there.&lt;/p&gt;
&lt;p&gt;By quantifying and understanding your app’s behavior more completely, you can then define desired behavior. This lets you detect when there are changes from that desired state. Further, you need to understand what causes the changes when they do happen. And then be able to resolve issues as they arise or adaptively adjust your definition of what’s acceptable for desired behavior.&lt;/p&gt;
&lt;p&gt;Finally, you need to be able to continuously improve your AI app, without degrading behavior. Teams can leverage the same adaptive tests to confidently roll out updates, swap in components, or add new features, without needing to take the app offline or kicking off lengthy new research and development cycles.&lt;/p&gt;
&lt;p&gt;This adaptive behavioral testing is what gives enterprise AI teams the coverage, clarity, and continuity necessary for productionizing AI apps and keeping them in production.&lt;/p&gt;
&lt;h3 id=&quot;productionize-faster-and-maximize-ai-uptime&quot;&gt;&lt;a href=&quot;#productionize-faster-and-maximize-ai-uptime&quot;&gt;&lt;strong&gt;Productionize faster and maximize AI uptime&lt;/strong&gt;&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;By having a continued understanding of consistent and reliable behavior, teams are able to build the confidence they need to ship higher value products faster, while minimizing risk to the business. Teams no longer need to build in a vacuum, and are able to account for the uncertainty of production usage while constantly adapting applications incrementally over time. Ultimately resulting in fewer production surprises and the ability to catch gradual shifts before users do. Plus, with a shared comprehensive view of desired behavior, the silos between development to production break down, resulting in faster and more predictable updates.&lt;/p&gt;
&lt;p&gt;What does this mean for you? You now have the bandwidth to tackle higher value AI use cases by leveraging the same repeatable and measurable processes to mitigate risk. Your business gets more impactful applications and your team can ship faster with confidence.&lt;/p&gt;
&lt;p&gt;Want to gain confidence in your AI applications? &lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Reach out&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; to Distributional to learn more.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>What is Distributional? Explaining AI behavior through the language of statistical distributions</title><link>https://distributional.com/blog/what-is-distributional-explaining-ai-behavior-through-the-language-of-statistical-distributions</link><guid isPermaLink="true">https://distributional.com/blog/what-is-distributional-explaining-ai-behavior-through-the-language-of-statistical-distributions</guid><description>Distributions as the native language of GenAI behavior: the AI Confidence Gap and statistical testing over behavior distributions. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 26 Feb 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;&lt;strong&gt;Statistical distributions&lt;/strong&gt; give us a language to understand the uncertain world around us by assigning likelihoods to future events of interest through repeated observation. In some ways, distributions are the lingua franca of chance—mutually intelligible across all domains—from describing the rate of faulty computer chips in manufacturing to estimating the size and frequency of waves in oceanography. Nowhere, however, is the parlance of chance more apropos than the domain of AI systems.&lt;/p&gt;
&lt;p&gt;In this piece I’ll discuss the opportunity for enterprises to increase the adoption of production AI systems—particularly GenAI systems—by gaining confidence that these systems will behave as desired once deployed. The key to this is evolving the notion of application testing so it is capable of communicating with AI in its native language of probability by way of statistical tests. By doing so, enterprises can deploy and maintain AI systems that continually pass these tests, signifying they are systems that have the highest tendency towards consistent desired behavior.&lt;/p&gt;
&lt;h3 id=&quot;ai-is-distributional&quot;&gt;&lt;a href=&quot;#ai-is-distributional&quot;&gt;AI is distributional&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Abstractly, &lt;a href=&quot;https://arxiv.org/pdf/2307.06435&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;GenAI systems are statistical distributions&lt;/a&gt; over natural language such that the goal is to predict the next token (letter, word, sentence, etc.) given a prompt and training corpus (which can be as large as the internet itself). From this premise, we get powerful applications like chatbots, LLM summarization, and AI agents. In building these systems we want to evaluate their distributions over performance metrics to find their propensity to correctly predict the next token or perform a specific task. Beyond that, we are interested in distributions which describe second-order properties (or metrics) of the system’s behavior itself, such as the likelihood that an LLM summarization system will deliver toxic, incoherent, or unreasonably long outputs. Because the outputs of AI systems can be represented by distributions over many metrics, their behavior can be described as “&lt;strong&gt;distributional&lt;/strong&gt;.”&lt;/p&gt;
&lt;p&gt;When it comes to developing and productionalizing these distributional AI systems, they require &lt;a href=&quot;https://distributional.com/blog/the-ai-software-development-lifecycle-a-practical-framework-for-modern-ai-systems&quot;&gt;different considerations&lt;/a&gt; than those of traditional software applications. For one, traditional software applications exhibit consistent and repeatable workflows without deviation, and therefore one can expect the same behavior from the same input. Ergo, my word processor application returns the character I desire to type with each keystroke, without fail.&lt;/p&gt;
&lt;p&gt;On the other hand, modern AI applications (both models and components) offer changing behavior over time for the same prompt or stimulus. Consider an LLM summarization application, for instance one used to summarize medical journals as part of an application sold to health practitioners. By its nature, this system may produce varying synopses of the same documents of interest, a property known as non-determinism. The model itself may also change over time, a quality known as non-stationarity. In the worst case scenario the outputs are inconsistent, inaccurate, too verbose, or even contain undesired words, thereby detracting from business value and causing operational risk to both vendor and consumer. In the best case scenario, all variations of summaries produced by the LLM are equally acceptable, accurate, and operable by the end user. This type of consistency is the desirable outcome for the vendor looking to increase revenue by providing such a beneficial AI application.&lt;/p&gt;
&lt;h3 id=&quot;the-ai-confidence-gap&quot;&gt;&lt;a href=&quot;#the-ai-confidence-gap&quot;&gt;The AI Confidence Gap&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Therefore, for the enterprise, the outcomes or behaviors of &lt;strong&gt;distributional&lt;/strong&gt; AI systems must align with business goals, providing the highest chance of desired behavior and lowest risk of undesired behavior. Today, there are many groups of tools on the market to help enterprises develop and deploy such distributional systems—mainly those focused on evaluation when building an AI application, and those focused on monitoring once that application is in production. Where these tools fall short is in providing value between development and production, a phenomenon known as the &lt;strong&gt;AI Confidence Gap&lt;/strong&gt;. That is, AI product teams (and the AI governance and compliance teams that oversee them), lack the confidence that the AI application created during development will behave as intended the minute real users utilize the deployed application at scale.&lt;/p&gt;
&lt;p&gt;This disconnect is so severe that it prevents AI production teams from deploying AI applications due to fear of operational risks. For this reason it is known market-wide that &lt;a href=&quot;https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;nearly 1 in 3 GenAI projects&lt;/a&gt; will be abandoned after proof of concept by the end of 2025. Whereas in traditional software, the confidence gap between development and production is closed by rigorous testing, given AI systems are governed by the laws of probability, a new concept of AI specific testing is required to fill the void.&lt;/p&gt;
&lt;h3 id=&quot;what-is-distributional&quot;&gt;&lt;a href=&quot;#what-is-distributional&quot;&gt;What is Distributional?&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Enter &lt;a href=&quot;https://distributional.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Distributional&lt;/a&gt;, the platform and company that bear the name of AI’s intrinsic essence, and introduced the concept of &lt;strong&gt;distributional testing for AI applications&lt;/strong&gt;. The key insight is that AI, by its nature, is a mixture of statistical distributions that can be queried to discover their central tendencies (and also their rare events). Hence, Distributional acts as an interpreter for AI, allowing product teams to statistically define consistency of AI application behavior, and then continuously confirm or identify changes in this behavior over time. Armed with statistical visibility into AI application behavior, AI product teams can confidently bridge the AI Confidence Gap.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/what-is-distributional-explaining-ai-behavior-through-the-language-of-statistical-distributions/67fe802c2b15ec95601db0e1_Kembey_Blog_Screenshot-1-2048x1148.avif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;The Distributional automated workflow for AI testing&lt;/p&gt;
&lt;h3 id=&quot;how-distributional-works&quot;&gt;&lt;a href=&quot;#how-distributional-works&quot;&gt;How Distributional works&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Functionally, Distributional provides an innovative approach that enables enterprise AI teams to understand, manage, and test the reliability of their AI systems so they can gain confidence in their behavior. Here is a simplified version of how this automated workflow works:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Integrate:&lt;/strong&gt; Point Distributional at your data storage for an AI application, any input, output, or intermediate data that the application produces&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Run:&lt;/strong&gt; Collect data on regular basis from this data storage using your current orchestration or CI tooling&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Derive:&lt;/strong&gt; Use Distributional’s eval library to derive testable behavioral properties from your data&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Test:&lt;/strong&gt; Apply statistical test templates to these properties that assess consistency in AI application behavior across distributions of these testable properties run over run&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Triage:&lt;/strong&gt; Analyze test results in the Distributional dashboard to decide whether the tests need to be calibrated to fit the app or app debugged to fit the tests&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resolve:&lt;/strong&gt; Recalibrate tests within Distributional or debug the application off platform with insights from Distributional&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This allows for apples to apples comparisons of the same production application over time or between new applications in development and those currently in production. As a result, teams feel confident in steady state AI app behavior and can quickly get back to steady state as their apps evolve.&lt;/p&gt;
&lt;h3 id=&quot;the-distributional-fingerprint&quot;&gt;&lt;a href=&quot;#the-distributional-fingerprint&quot;&gt;The Distributional Fingerprint&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;In the next piece in this series I will go deeper into this workflow. Although Distributional can be applied to any AI/ML application I will make this concrete by showing you how it works on a simple Q&amp;amp;A RAG application. In doing so we will see Distributional’s main features, namely the ability to produce an AI application’s &lt;strong&gt;Distributional Fingerprint&lt;/strong&gt; (i.e. its unique baseline mixture of characteristic distributions), automate and apply statistical tests, detect and understand behavioral change to the &lt;strong&gt;Distributional Fingerprint&lt;/strong&gt;, and enable AI governance through a standardized, repeatable, consistent, and visible testing process.&lt;/p&gt;
&lt;p&gt;Interested in trying Distributional? Here are some next steps and ways to learn more:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Get access to our product&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;Reach out to our team&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://distributional.com/blog&quot;&gt;Read more on our blog&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded><category>archive</category></item><item><title>The rise of AI platforms</title><link>https://distributional.com/blog/the-rise-of-ai-platforms</link><guid isPermaLink="true">https://distributional.com/blog/the-rise-of-ai-platforms</guid><description>Why internal AI platforms echo the 2018 ML platform wave: the emerging component stack and how platform teams sequence adoption. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 19 Feb 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;“This is the most SigOpt-like solution ever,” said the AI executive of a Global 2000 energy company, referring to the &lt;a href=&quot;https://sigopt.org/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;AI optimization startup&lt;/a&gt; we built and sold to Intel in 2020. “You’ve started with the hardest technical problem in the GenAI stack and haven’t built out the other components I need,” he continued with a grin.&lt;/p&gt;
&lt;p&gt;He proceeded to explain the components he had already built or planned to build for his internal AI platform. He had started out by hosting an LLM and releasing a simple version of it to a few teams to see what use cases were most prominent. Now that he had this initial information, he felt confident in standardizing his AI platform. “Not only do I feel confident in my AI platform now, but we need it. I can’t responsibly scale up usage anymore unless we put a platform in place.”&lt;/p&gt;
&lt;p&gt;This is similar to dozens of conversations I’ve had since. Every team I talk to is building, has plans to build, or has already built an AI platform. In this post, I’ll provide a snackable summary of what I’ve learned so far about the latest platform trend.&lt;/p&gt;
&lt;h3 id=&quot;ml-ai-platforms-in-2018-2025&quot;&gt;&lt;a href=&quot;#ml-ai-platforms-in-2018-2025&quot;&gt;ML AI platforms in 2018 2025&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;If this all sounds familiar, it should. In 2018, I sat down with Sam Charrington, founder of &lt;a href=&quot;https://twimlai.com/resources/the-definitive-guide-to-machine-learning-platforms/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;TWIML&lt;/a&gt;. Earlier that year, I had joined Scott Clark and the SigOpt team. “Every customer I talk to discusses integrating us with their ML platform,” I explained to Sam. He nodded along with a wry smile suggesting I was explaining something he already knew. Six months later, we sponsored his independent research that led to the ML Platforms &lt;a href=&quot;https://twimlai.com/resources/the-definitive-guide-to-machine-learning-platforms/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;guide&lt;/a&gt;, podcast &lt;a href=&quot;https://twimlai.com/podcast/twimlai/supporting-rapid-model-development-two-sigma-scott-clark-matthew-adereth/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;series&lt;/a&gt;, and &lt;a href=&quot;https://twimlai.com/conf/twimlcon/2022/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;event&lt;/a&gt;. Six months after this, &lt;a href=&quot;https://en.wikipedia.org/wiki/MLOps&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;MLOps&lt;/a&gt; became an industry standard category describing this toolchain. That won’t be the last branding moment I miss.&lt;/p&gt;
&lt;p&gt;You won’t be surprised that the critical elements in this stack focused on feature engineering, training, and serving. Feature stores, experiment tracking, orchestration, and monitoring were a few of the relatively stable categories that emerged and seeded incumbent startup winners. All of them were designed with training models as the critical workflow, and most had structured data in mind.&lt;/p&gt;
&lt;p&gt;Once trained, these models were relatively stable. I had a recent partner tell me that when they were building XGBoost models for financial services companies, they’d retrain them on an annual basis at most. Generative AI broke this model.&lt;/p&gt;
&lt;h3 id=&quot;emerging-ai-platform-components&quot;&gt;&lt;a href=&quot;#emerging-ai-platform-components&quot;&gt;Emerging AI platform components&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Pre-trained generative AI models have taken industry by storm. Most obviously, these models don’t need training, which obviates a good chunk of the ML platform stack. These models just work, and tend to work well.&lt;/p&gt;
&lt;p&gt;There has also been increasing standardization of components, dependencies, packages, and tooling that make it easier to create complex chains of AI components—agents or otherwise. AI today is less experimental, more engineered.&lt;/p&gt;
&lt;p&gt;So what does this mean for AI platforms? GenAI is in high demand. These models are powerful, but unpredictable. They are also expensive to run. They can be hard to get to reliable performance, and even harder to understand. Most teams I’ve talked to are cobbling together some variety of these components:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Model and/or compute management&lt;/strong&gt;, including token cost optimization or prioritization&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Proxy&lt;/strong&gt; to seamlessly plug in any LLM to an app with various forms of &lt;strong&gt;routing&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Controls and permissions&lt;/strong&gt; by group, team, sometimes even individual&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Orchestration&lt;/strong&gt;, including hosting, inference, optimization, chaining&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Observability&lt;/strong&gt;, including traces, logs, metrics, evals&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Testing&lt;/strong&gt;, including consistency across all data types from all components and properties&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;User interface&lt;/strong&gt; that creates a standard way for a user to access an LLM&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evals, prompt playgrounds, and other dev tools&lt;/strong&gt; to iterate on application performance&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Guardrails, firewalls,&lt;/strong&gt; and other systems to pre-empt bad behavior or attacks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I’m certainly missing many components, and perhaps differently characterizing how you may think of some of these. One hero theme here, however, is that the goal of these platforms is to accelerate and scale generative AI productionalization—and ensure companies are allocating resources in the right ways. If done well, these platforms will result in greater return on AI investment.&lt;/p&gt;
&lt;p&gt;This is a new stack. Even if they have ambitions of marrying the two together at various points, teams aren’t building off of their ML platforms. They are building new AI platforms to enable this workflow. And teams that get this right will be able to extract business value out of GenAI faster than the competition.&lt;/p&gt;
&lt;h3 id=&quot;the-case-for-best-in-class&quot;&gt;&lt;a href=&quot;#the-case-for-best-in-class&quot;&gt;The case for best in class&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;So how do you get started? This space is moving so fast. Seemingly every day there’s a new entrant with a new take on one of these tools. This can be intimidating, confusing, and hard to navigate. Amidst this chaos, it can be tempting to try to use an end-to-end stack that has 50% of these components and does them at 80%.&lt;/p&gt;
&lt;p&gt;This is fool’s gold. You start faster, but quickly run into problems as you scale. What started out easy, becomes incredibly hard as you shift resources from building new apps to debugging the underlying platform—and layers of APIs that support it. The teams I’ve seen moving fastest on generative AI experiment quickly to define what works and what doesn’t. But then they shift to a more reliable, scalable, and controllable stack as they scale. This approach gives them a few benefits:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Build &lt;em&gt;and&lt;/em&gt; buy&lt;/strong&gt;: Taking this approach makes it easier to prioritize which components your team will build as part of your core stack that you control, and which you will buy. And puts you in a position to build MVP versions of these components, and upgrade them with vendor solutions as the market becomes clearer with winners and losers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Integrate without lock-in&lt;/strong&gt;: This approach takes a bit more investment, but gives you these high-end capabilities without a threat of lock-in. A leading vendor today may be gone tomorrow. New innovation in the underlying models may result in new needs on the platform side. This approach also allows you to build on the prior work of your team, with the ability to plug in outputs from past solutions as the space evolves.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Standardization &lt;em&gt;and&lt;/em&gt; customization&lt;/strong&gt;: There isn’t a template for what these platforms should look like. You need to standardize how this is done internally, but your needs may be different than someone else. Building it up yourself with best-in-class components gives you the chance to have the best of both of these worlds—standardization and customization.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Put together, this suggests a best-in-class approach that doesn’t lock you into something that doesn’t scale or evolve with your needs over time.&lt;/p&gt;
&lt;h3 id=&quot;how-to-get-started&quot;&gt;&lt;a href=&quot;#how-to-get-started&quot;&gt;How to get started&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Inertia can be hard to overcome. Our own team has experienced short-term paralysis looking at various options for our own internal GenAI efforts. The best advice I have is to get started as fast as possible and enable the usage of at least a single GenAI model for various applications, whether an API endpoint or self-hosted. Maybe take an existing NLP component and upgrade it with an LLM. These can give you good initial data on potential usage. And once you have underlying app data, you can begin to build out an understanding of how these apps will work.&lt;/p&gt;
&lt;p&gt;That’s where we come in. In his recent TWIML &lt;a href=&quot;https://twimlai.com/podcast/twimlai/evolving-mlops-platforms-for-generative-ai-and-agents/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;interview&lt;/a&gt;, Capital One AVP of AI/ML Enterprise Platforms Abhijit Bose says, “Observability…becomes very, not just important, but also very complex in the LLM world.” I agree.&lt;/p&gt;
&lt;p&gt;It is never too early to get testing infrastructure in place—start small and scale it up as you grow. &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;Reach out&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; to start the conversation with us today.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Driving confidence in AI: Foundational distributional testing explained through a RAG example</title><link>https://distributional.com/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example</link><guid isPermaLink="true">https://distributional.com/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example</guid><description>Foundational distributional testing walked through a RAG example: baselines, distributional comparison, experiments, and production data. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Wed, 29 Jan 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;In this article, I’d like to share some examples to help readers gain intuition around what it means to statistically test a RAG app, using the example of a knowledge base on every engineer’s favorite topic, cats. The goal is to help develop some intuition around the type of data collected when running a GenAI app, and how this can be helpful when understanding the behavior of the app. In this example, we are looking at a situation with changing user behavior, but we can easily imagine a scenario where it’s one of the components or the RAG itself that is causing problems.&lt;/p&gt;
&lt;h3 id=&quot;looking-at-a-rag-app-as-black-box&quot;&gt;&lt;a href=&quot;#looking-at-a-rag-app-as-black-box&quot;&gt;Looking at a RAG app as black box&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Let’s start with how users see a RAG app. On the surface, a RAG app is just a black box—a user asks a question, and the app gives back an answer.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example/67ff3122664e237940e37547_image6.jpg&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;When we ask a question: “Why do cats meow?”&lt;/p&gt;
&lt;p&gt;We get an answer back: “Cats typically meow primarily as a form of communication or attention-seeking, rather than due to hunger or need for food. Meowing can indicate pleasure, excitement, or affection, often when a human is nearby or trying to interact with the cat.”&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[{&amp;#39;Question&amp;#39;: &amp;#39;Why do cats meow?&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Answer&amp;#39;: &amp;#39;Cats typically meow primarily as a form of communication or&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       attention-seeking, rather than due to hunger or need for food.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       Meowing can indicate pleasure, excitement, or affection, often&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       when a human is nearby or trying to interact with the cat.&amp;#39;}]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Whether or not we were actually able to answer the question correctly is not the topic of this post. For now, we are interested in knowing whether the question is being answered using information for our knowledge base, or if our app is simply just using the large language model to answer the question.&lt;/p&gt;
&lt;h3 id=&quot;looking-at-the-logs-for-black-box-app&quot;&gt;&lt;a href=&quot;#looking-at-the-logs-for-black-box-app&quot;&gt;Looking at the logs for black box app&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Now, the problem is that we have no idea what’s happening within this black box. On its own, the question-answer pair doesn’t tell us if the answer comes from knowledge that is embedded within the large language model, or if it comes from the knowledge base connected to our application. Is the answer a hallucination—or something we can trust?&lt;/p&gt;
&lt;p&gt;Just looking at the question-answer pair doesn’t give us enough information. If we want to understand how our RAG app got the answer, we’ll need to look under the hood.&lt;/p&gt;
&lt;h3 id=&quot;splitting-your-rag-app-into-components&quot;&gt;&lt;a href=&quot;#splitting-your-rag-app-into-components&quot;&gt;Splitting your RAG app into components&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;To start, we need to break down our RAG app into components. Let’s assume we have a basic vanilla RAG application that relies on a single information database.&lt;/p&gt;
&lt;p&gt;Starting with a question, we look for similar contexts in our vector database and identify a number of documents which hopefully contain the answer to that question. We then pair those documents with the actual question and give it to our LLM, which generates an answer. For this example we’ll identify three documents every time we generate and answer.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example/67ff3137b418d2f9586006e2_image1.jpg&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;All of this is controlled by our instruction prompt, where we tell our application how to behave and what to do.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;instruction_prompt_context = &amp;quot;&amp;quot;&amp;quot;     You are a helpful chatbot designed to provide answers based on the     retrieved documents.         Follow these guidelines:         1. You can only provide responses less than 30 words.         2. Focus on providing insights directly from the retrieved            documents and avoid speculative or unsupported answers.         3. Maintain a professional tone suitable for executives and            avoid technical jargon unless explicitly requested.      Always ensure the output is accurate, relevant, and adheres to these     instructions.      These are the retrieved documents:     &amp;quot;&amp;quot;&amp;quot;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;One of the things that we are telling our application is:&lt;/p&gt;
&lt;p&gt;&lt;code&gt;2. Focus on providing insights directly from the retrieved documents and avoid speculative or unsupported answers.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;How to test whether or not our application is actually following this guideline is the focus for this blogpost. How to test for instruction 1. is pretty self explanatory (simply check how many words are in the response), and how to test for 3. is something that we will cover in a later blogpost.&lt;/p&gt;
&lt;h3 id=&quot;looking-at-the-logs-for-component-app&quot;&gt;&lt;a href=&quot;#looking-at-the-logs-for-component-app&quot;&gt;Looking at the logs for component app&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Now, when we look at the log for a single app usage, we still have the question-answer pair but now we also have a list of all the documents that have been retrieved, along with similarity scores for each of them. For this particular example, the similarity scores are computed using the same embedding model as for the knowledge base and the cosign similarity score.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[{&amp;#39;Question&amp;#39;: &amp;#39;Why do cats meow?&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Answer&amp;#39;: &amp;#39;Cats typically meow primarily as a form of communication or&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       attention-seeking, rather than due to hunger or need for food.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       Meowing can indicate pleasure, excitement, or affection, often&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       when a human is nearby or trying to interact with the cat.&amp;#39;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_1_text: &amp;#39;A cat almost never meows at another cat, mostly just humans.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                Cats typically will spit, purr, and hiss at other cats&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_1_QD_sim: &amp;#39;0.72&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_1_DA_sim: &amp;#39;0.73&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_2_text: ...}]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;These similarity scores measure the similarities between the question and different pieces of data within the vector database [QD_sim]. This tells us whether or not the information exists within our database.&lt;/p&gt;
&lt;p&gt;Another thing we can compute is the similarity between the document that I’ve retrieved and the answer that was generated [DA_sim]. This tells us whether or not the LLM is actually using the information provided to generate an answer. Together, these give us more insight into where the information used to create the answer is coming from.&lt;/p&gt;
&lt;h3 id=&quot;creating-a-baseline&quot;&gt;&lt;a href=&quot;#creating-a-baseline&quot;&gt;Creating a baseline&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;To better understand our RAG app’s behavior through a statistical lens, we need to create a baseline. A baseline is a number of logs that answer questions in the way that we want them to. How we are coming up with our baseline is not something that we are going to cover in this blogpost.&lt;/p&gt;
&lt;p&gt;We can create a baseline using real data from our customers, or it can be from a set of questions we generated ourselves—having a sense of what questions should look like and the kind of behavior we want to see.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[{&amp;#39;Question&amp;#39;: &amp;#39;Why do cats meow?&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Answer&amp;#39;: &amp;#39;Cats typically meow primarily as a form of communication or&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       attention-seeking, rather than due to hunger or need for food.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       Meowing can indicate pleasure, excitement, or affection, often&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       when a human is nearby or trying to interact with the cat.&amp;#39;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_1_text: &amp;#39;A cat almost never meows at another cat, mostly just humans.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                Cats typically will spit, purr, and hiss at other cats&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_1_QD_sim: &amp;#39;0.72&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_1_DA_sim: &amp;#39;0.73&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_2_text: ...},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;{&amp;#39;Question&amp;#39;: &amp;#39;Can cats get sunburned?&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Answer&amp;#39;: &amp;#39;Yes, cats can indeed get sunburned. Frequent exposure to the sun,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;             especially on their white fur and exposed skin areas like ears,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;             can cause skin damage leading to sunburn.&amp;#39;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_1_text: &amp;#39;Cats with white fur and skin on their ears are very prone to&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                sunburn. Frequent sunburns can lead to skin cancer. Many white&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                cats need surgery to remove all or part of a cancerous ear.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                Preventive measures include sunscreen, or better, keeping the&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                cat indoors.&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_1_QD_sim: &amp;#39;0.82&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_1_DA_sim: &amp;#39;0.88&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Doc_2_text: ...},]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then we take that baseline and create distributions from the baseline. What we see in the histograms below is what we call a “run” in Distributional, which are the similarity scores for our baseline over a set period of time, such as an hour or a day.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example/67ff316db1a6c242aa4f2579_48cf5d66-266f-4906-8beb-22e7783479a9.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;h3 id=&quot;looking-at-a-single-result&quot;&gt;&lt;a href=&quot;#looking-at-a-single-result&quot;&gt;Looking at a single result&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Let’s take a look at a single question-answer pair within these distributions, going back to our original question “Why do cats meow?”&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[{&amp;#39;Question&amp;#39;: &amp;#39;Why do cats meow?&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Answer&amp;#39;: &amp;#39;Cats typically meow primarily as a form of communication or&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       attention-seeking, rather than due to hunger or need for food.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       Meowing can indicate pleasure, excitement, or affection, often&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       when a human is nearby or trying to interact with the cat.&amp;#39;}]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The answer seems right to me, but let’s look at the text retrieved from the documents ranked highest for similarity.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example/67ff319bb47f03aab65d1223_36f443ba-6ff5-435b-9549-84dcbe126b35-1.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;These documents look good to me, but let’s introduce an experiment to test this and gain more intuition on what things look like when they go off the rails.&lt;/p&gt;
&lt;h3 id=&quot;introducing-an-experiment&quot;&gt;&lt;a href=&quot;#introducing-an-experiment&quot;&gt;Introducing an experiment&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;To get an understanding of how it looks, when things are not going as planned, we can introduce a simple experiment, and look at how this differs from our baseline. To do that, we are asking our app a whole bunch of questions about other animals than cats. Similarly, to the baseline we then log the similarity scores and plot those in a histogram (red) together with our baseline (blue).&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example/67ff31b6664e237940e40650_61ee45cb-3f4d-43d7-a468-600afceea77b.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Here we see that pretty much all of the similarity scores are significantly lower for the animal-questions than the cat-questions, meaning that it is less likely that our RAG-app has found the correct information in our knowledge base to answer the question.&lt;/p&gt;
&lt;h3 id=&quot;looking-at-single-result&quot;&gt;&lt;a href=&quot;#looking-at-single-result&quot;&gt;Looking at single result&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Let’s take a look at a single log to better understand what’s happening. If we ask a question about fish, the answer seems accurate.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[{&amp;#39;Question&amp;#39;: &amp;#39;Question: What role do cleaner fish play in maintaining the&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;               health of other fish?&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Answer&amp;#39;: &amp;#39;Answer: Cleaner fish help maintain the health of other fish by&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;             removing parasites and organic waste from their own bodies. This&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;             process is known as biofertilization or biological control.&amp;#39;}]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;But when we look at our distributions and the text from our similar documents, it’s clear that this answer is not coming from within our database. All of the text referenced is about cats, not fish, and the question-answer pair does not align with our baseline distributions either.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example/67ff31dbf102567240fd5028_51b8fb63-6eea-4511-b7a5-1bec0e4bbc2b.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;So now, instead of having a human go though all of our production logs and categorize whether or not the LLM is hallucinating, we can use our simple to compute similarity scores the get a good understanding of whether or not that is the case, and significantly reduce the need for humans in the loop.&lt;/p&gt;
&lt;h3 id=&quot;introducing-production-data&quot;&gt;&lt;a href=&quot;#introducing-production-data&quot;&gt;Introducing production data&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Let’s continue to explore how our app responds when we introduce real production data. Although we hope that our users will only ask cat questions, they may start asking questions about other topics. For example, octopuses.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8;overflow-x:auto&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[{&amp;#39;Question&amp;#39;: &amp;#39;Question: How do octopuses camouflage themselves in their&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;               environment?&amp;#39;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#39;Answer&amp;#39;: &amp;#39;Answer: Octopuses use color-changing cells called&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;             chromatophores to camouflage themselves, making it difficult for&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;             predators to detect them. They also use pattern recognition and&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;             texture matching to blend in with their surroundings.&amp;#39;}]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;At a quick glance, we can quickly see which question-answer pairs are cause for concern. For the question “How do octopuses camouflage themselves in their environment?” we can see that the documents being retrieved have nothing to do with octopuses. Instead, this answer is being generated with knowledge that the LLM has.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/driving-confidence-in-ai-foundational-distributional-testing-explained-through-a-rag-example/67ff31fc4768bfcf2dced390_16c420df-8a7b-4af7-a352-348b4be48b14.png&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Even though I’ve specified that my app can only use information I’ve provided to generate an answer, it’s clearly going off the rails and generating what presumably is the correct answer using information embedded in the model. While this is likely not too big of a problem for an app focused on cats, it could be very bad for organizations focused on finance, HR, law, or healthcare. Here the stakes are much higher and we don’t want our RAG app to hallucinate answers. Therefore, it is important for companies to have an automated way, similar to what we just walked through, to detect these things at scale.&lt;/p&gt;
&lt;p&gt;To make the above insights actionable, we need to start thinking about why the change is occurring. In this particular example, it is pretty clear why things are starting to look different – our users are asking questions from an unintended domain – which can potentially be solved by changing the way our users interact with the app. However, not all changes are a function of our users, a lot of changes come from drift in models and continuous changes to the app. Here we need to start thinking about whether the problem comes from our instruction prompt, our vector database, retrieval algorithms or something completely different. This is what we do at Distributional, helping you understand where and when these changes are happening, and from there give you the tools you need to address the problems.&lt;/p&gt;
&lt;h3 id=&quot;automating-testing-for-rag-apps&quot;&gt;&lt;a href=&quot;#automating-testing-for-rag-apps&quot;&gt;Automating testing for RAG apps&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Tests like these are just a few of many that teams should run to continuously test their RAG apps.&lt;/p&gt;
&lt;p&gt;With Distributional, modeling teams can set up and automate statistical tests to help them understand how the performance of their GenAI applications has shifted over time and get alerts when things go off the rails. Our goal is to help teams develop confidence in the AI applications they’re taking to production, so they can rely on more than just a gut check to ensure their applications are working as intended.&lt;/p&gt;
&lt;p&gt;Interested in exploring how Distributional can help your organization? &lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Sign up&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; for access to our product or &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;get in touch&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; with our team to learn more.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>The AI Software Development Lifecycle: A practical framework for modern AI systems</title><link>https://distributional.com/blog/the-ai-software-development-lifecycle-a-practical-framework-for-modern-ai-systems</link><guid isPermaLink="true">https://distributional.com/blog/the-ai-software-development-lifecycle-a-practical-framework-for-modern-ai-systems</guid><description>The AI SDLC framework: how the development lifecycle morphs across traditional software, ML, and GenAI, with risk rising as determinism falls. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Thu, 09 Jan 2025 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;I recently found myself in a discussion with a group of ML engineers who were debating whether traditional ML development practices could transfer to GenAI applications. The conversation highlighted a crucial point: while there’s significant overlap between traditional ML and GenAI development lifecycles, the unique characteristics of GenAI systems demand a fresh perspective on how we build, deploy, and maintain these applications.&lt;/p&gt;
&lt;p&gt;The software development lifecycle (SDLC) is a structured approach to software development that guides development teams through the stages of designing, building, and deploying high-quality software. Its primary goal is to reduce project risks through proactive planning, ensuring that the software meets expectations both during production and throughout its lifecycle. Traditional software development follows a predictable, linear path with clear requirements and deterministic outcomes. Machine learning development shifts the focus to data collection and pre-processing, model training and validation with clear performance metrics, and model versioning. GenAI development introduces a new paradigm centered on prompt engineering and model fine-tuning, managing non-deterministic outputs, evaluating contextual understanding and response quality, and balancing cost with performance.&lt;/p&gt;
&lt;p&gt;While traditional software and ML applications rely on concrete metrics and testing, GenAI requires more flexible evaluation methods and continuous adaptation to model changes. &lt;strong&gt;The key distinction lies in how risk increases and certainty decreases as we move from traditional software (most certain) to ML (less certain) to GenAI (least certain), requiring increasingly adaptive development approaches.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-ai-software-development-lifecycle-a-practical-framework-for-modern-ai-systems/67ffe9583ce777b36ed4456a_Blog.jpg&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Types of software compared by risk and against the level of non-determinism, non-stationarity, and complexity.&lt;/p&gt;
&lt;p&gt;With over two decades of experience helping to build and drive the evolution from traditional software development to machine learning and now to generative AI, I have been at the forefront of these transformative technological shifts. This experience has shown me the critical importance of establishing a clear framework for the AI Software Development Lifecycle (AI SDLC). While this framework may be simplified, it captures the essential elements that teams need to consider when developing AI systems.  Below, I present a simplified version of the AI SDLC and a discussion of the different phases of the lifecycle, including common tooling that many teams use today.&lt;/p&gt;
&lt;h3 id=&quot;the-ai-sdlc-a-simplified-framework&quot;&gt;&lt;a href=&quot;#the-ai-sdlc-a-simplified-framework&quot;&gt;The AI SDLC: A simplified framework&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;There are four key phases in the simplified view of the AI SDLC: Explore, Build, Deploy, and Observe.  Though these phases might resemble the traditional SDLC, the tooling in AI is more complex, and the AI SDLC differs much more significantly at the Deploy and Observe stages due to the non-deterministic nature of GenAI applications.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://distributional.com/images/blog/the-ai-software-development-lifecycle-a-practical-framework-for-modern-ai-systems/67fe7f5285b78bceae3fb686_AI-SDLC-Chart.avif&quot; alt loading=&quot;lazy&quot; decoding=&quot;async&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Simplified AI SDLC Diagram: Explore, Build, Deploy, and Observe stages.&lt;/p&gt;
&lt;h4 id=&quot;explore&quot;&gt;&lt;a href=&quot;#explore&quot;&gt;Explore&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;The exploration phase involves defining the problem scope, understanding available data, and selecting appropriate model architectures. While this stage remains relatively consistent between ML and GenAI projects, GenAI applications might fast-track this phase, moving directly to building with pre-trained foundation models and prompt engineering. Teams typically leverage basic data analysis tools, IDEs, and visualization frameworks, with GenAI projects potentially incorporating prompt playgrounds for initial experimentation. If you don’t spend enough time understanding and defining the problem, you risk solving the wrong problem. A goal of the explore phase is to gather enough information such that you set yourself up for success at the subsequent phases.&lt;/p&gt;
&lt;h4 id=&quot;build&quot;&gt;&lt;a href=&quot;#build&quot;&gt;Build&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;This stage focuses on rapid prototyping and offline evaluation to validate the solution feasibility. The build phase marks a significant divergence between ML and GenAI approaches. Traditional ML emphasizes model training with well-defined performance metrics, while GenAI development centers on iterative prompt engineering and semantic analysis of responses, often with less concrete success metrics.&lt;/p&gt;
&lt;p&gt;The development tooling landscape has exploded, with numerous startups offering solutions that offer everything from “prompt playgrounds” and debugging to log viewers and dataset annotation. While some frameworks provide extensive functionality and a polished interface, teams can succeed with a minimal toolset focused on prompt management and basic testing. While this phase is intentionally fluid and experimental, testing during the Build phase typically serves a few different purposes: proactively addressing anticipated deployment constraints, recreating production issues in a development environment for debugging, or working to drive improved performance in responses. Outside of these scenarios, most systematic testing is better suited for the Deploy phase. While the Build phase may appear less structured than traditional software development, its dynamic and exploratory nature is purposeful—enabling teams to rapidly test hypotheses and discover optimal solutions.&lt;/p&gt;
&lt;p&gt;Because of this open-ended nature, it is common for AI applications to get stuck at the Build phase—in particular, during the first iteration around the SDLC cycle. If you are able to explicitly define what success looks like before you begin building, that can help provide the confidence you need to feel ready to move on to the next phase.&lt;/p&gt;
&lt;h4 id=&quot;deploy&quot;&gt;&lt;a href=&quot;#deploy&quot;&gt;Deploy&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Pre-production deployment involves ensuring application scalability, generalization capability, and stability. While both ML and GenAI systems face similar challenges with multi-component deployments, GenAI applications introduce additional complexity through their inherent non-determinism and potential dependency on third-party LLMs. This necessitates more rigorous consistency checking and validation procedures. Currently, there’s a noticeable gap in specialized deployment tools for AI systems, making testing frameworks crucial for bridging development and production environments.&lt;/p&gt;
&lt;p&gt;AI applications are typically multi-component systems, so it is important at the Deploy phase to carefully test where the components fit together in a production-simulated environment. Unexpected outputs out of an LLM, for example, can easily cascade across the system and create instability within other components, potentially causing a system failure.&lt;/p&gt;
&lt;h4 id=&quot;observe&quot;&gt;&lt;a href=&quot;#observe&quot;&gt;Observe&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;The observation phase encompasses monitoring and maintenance of deployed AI systems. While ML and GenAI share common monitoring needs, GenAI presents unique challenges:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Optimization becomes critical due to reference model costs and a lack of control. For example, if there’s a shift in the distribution of your inputs that causes a large increase in the number of tokens returned by an LLM, you could incur unexpected expenses.&lt;/li&gt;
&lt;li&gt;Unstructured data complicates monitoring and testing. For example, a customer service chatbot needs to be able to recognize misspellings and unexpected input that might not be present in the training or offline testing data.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Non-deterministic behavior requires statistical analysis of metric distributions in order to understand whether something meaningful has changed. For example, if the distribution shifted from unimodal to bimodal, you wouldn’t pick that up looking at averages.&lt;/p&gt;
&lt;p&gt;The tooling ecosystem includes inference optimization, routing systems, monitoring platforms, and guardrail implementations. Testing frameworks incorporate and complement these tools by providing continuous validation of system behavior. Without observability, the AI SDLC would be a linear process ending at the deployment step. However, the insights gained at the Observe phase allow us to circle back around, adjusting any assumptions that turned out not to be correct, and improving the models at the Build stage to more closely meet desired expectations and goals.&lt;/p&gt;
&lt;h3 id=&quot;moving-through-the-lifecycle&quot;&gt;&lt;a href=&quot;#moving-through-the-lifecycle&quot;&gt;Moving through the lifecycle&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;The AI SDLC differs from the traditional SDLC is that teams typically progress through these stages in cycles, rather than linearly. A GenAI project might start with rapid prototyping in the Build phase, circle back to Explore for data analysis, then iterate between Build and Deploy as the application matures. The Observe phase provides continuous feedback that often triggers new cycles through earlier stages.&lt;/p&gt;
&lt;p&gt;The iterative nature of the AI SDLC provides a mechanism by which the application can be continuously improved based on insights gleaned in the Observe phase. Without continuous feedback from the Observe phase, there’s a substantial risk that the application does not behave as expected—for example, the output of an LLM could provide nonsensical or incorrect responses.&lt;/p&gt;
&lt;p&gt;Additionally, the performance of the application degrades over time as the distributions of the input data shifts.&lt;/p&gt;
&lt;h3 id=&quot;looking-ahead&quot;&gt;&lt;a href=&quot;#looking-ahead&quot;&gt;Looking ahead&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;The AI SDLC continues to evolve as we gain more experience with GenAI systems. As higher complexity applications of AI (e.g. RAG, agents) get created, the SDLC will evolve along with that. While this simplified framework provides a foundation for development practices, teams should adapt it to their specific needs and constraints.&lt;/p&gt;
&lt;h4 id=&quot;about-distributional&quot;&gt;&lt;a href=&quot;#about-distributional&quot;&gt;About Distributional&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Distributional is building the platform for AI testing, to help all AI teams gain confidence in the reliability of their applications, no matter how the SDLC evolves. &lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Sign up&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; to get access.&lt;/p&gt;</content:encoded><category>archive</category></item><item><title>Why testing is an enterprise problem that requires an enterprise solution</title><link>https://distributional.com/blog/why-testing-is-an-enterprise-problem-that-requires-an-enterprise-solution</link><guid isPermaLink="true">https://distributional.com/blog/why-testing-is-an-enterprise-problem-that-requires-an-enterprise-solution</guid><description>Part 3 of the enterprise AI testing series: why production-grade AI testing needs an enterprise platform, not developer-tool latitude. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Thu, 31 Oct 2024 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;What do AI leaders &lt;em&gt;really&lt;/em&gt; think about AI today? We’ve had over 1,000 hours of conversations with AI leaders at more than 100 Fortune 500 enterprises. This is the third post in a three-part series summarizing lessons from these conversations, focused on why AI testing is an enterprise problem that requires an enterprise solution. Read the first two posts &lt;a href=&quot;https://distributional.com/blog/why-generative-ai-creates-unique-testing-challenges&quot;&gt;here&lt;/a&gt; and &lt;a href=&quot;https://distributional.com/blog/why-all-ml-and-ai-use-cases-need-standardized-testing&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Before jumping in, it’s worth noting that I’m the CRO here at Distributional and the only one of our 25 team members without a technical degree (we have more than 35 technical degrees across 25 team members). My take is therefore intended to be more accessible, less technical and, hopefully, a quick read. If you want to dive deeper, &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;let’s find time to talk&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;
&lt;h3 id=&quot;enterprise-versus-developer-products&quot;&gt;&lt;a href=&quot;#enterprise-versus-developer-products&quot;&gt;Enterprise versus developer products&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Even the most established enterprises typically give a lot of latitude to AI development teams in tool selection as they’re iterating or prototyping in the model and application development phase. This includes evals, where teams typically start from open source benchmarks and evolve them with their own domain expertise to define and get to a level of performance that fits their needs. The goal of these processes is to minimize friction or constraints in a relatively free-form research process.&lt;/p&gt;
&lt;p&gt;AI testing, however, requires a very different approach. The goal of testing is to enable teams to get confidence with a definition of steady state for any AI application, confirm that it still meets this definition, and, where it deviates, figure out what needs to evolve or be fixed to reach steady state once again. This process needs to be discoverable, logged, organized, consistent, integrated and scalable. It is an enterprise problem that needs an enterprise solution. This approach requires that you start standardizing tests as checks in development, evolve them into a standard suite in deployment, and expand on them to cover scenarios where multiple components are shifting in production.&lt;/p&gt;
&lt;p&gt;Below are a few more specific challenges related to AI testing that lend themselves to an enterprise solution.&lt;/p&gt;
&lt;h3 id=&quot;enterprise-testing-needs&quot;&gt;&lt;a href=&quot;#enterprise-testing-needs&quot;&gt;Enterprise testing needs&lt;/a&gt;&lt;/h3&gt;
&lt;h4 id=&quot;production-grade-consistency&quot;&gt;&lt;a href=&quot;#production-grade-consistency&quot;&gt;Production-grade consistency&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;There needs to be a standard way to collect testing data, augment this data, define tests, evaluate results, validate actions, and give visibility into the testing process in every phase of the AI/ML software lifecycle. And these tests need to study behavioral properties for all components for every AI application. This level of depth and continuity gives enterprises visibility into AI app behavior that translates to confidence in productionalizing these apps. And this process starts at the very beginning of the application lifecycle. As one ML platform product leader in financial services explained, “I need a set of pre-deployment checks that I run consistently to know what to expect with each AI app in production. But this process of defining the tests needs to start in development and evolve in production.”&lt;/p&gt;
&lt;h4 id=&quot;multi-component-visibility&quot;&gt;&lt;a href=&quot;#multi-component-visibility&quot;&gt;Multi-component visibility&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;It’s possible to use free or open source libraries to run one-off evals that give you point-in-time confidence of a given AI or ML model. But AI applications aren’t typically a single model—rather, they’re a series of components that may be constantly shifting in different directions. A single developer rarely controls the whole pipeline, so a developer tool is not a good solution for gaining visibility into what is actually wrong across the full pipeline. Instead, teams need an enterprise solution capable of standarding how tests are run for every component of an AI application so teams know where the issue is.&lt;/p&gt;
&lt;h4 id=&quot;multi-team-usability&quot;&gt;&lt;a href=&quot;#multi-team-usability&quot;&gt;Multi-team usability&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;There are multiple users with diverse needs in enterprise AI testing. AI development, engineering, product, and governance teams may all need to run tests, but have different workflow and goals associated with this testing process. Sometimes the goal is enabling complete configurability so an AI developer can customize the experience to their domain. Other times, it requires automating metric computation and test configuration so users can consume the results of these tests in a more abstract way. Any enterprise platform needs to enable this diverse set of users and use cases. Distributional is built to do this. An AI product leader at a large consumer technology company told us, “I love how your platform empowers me to explore my applications without needing to get in the weeds of selecting the right tests or even defining the full set of metrics.”&lt;/p&gt;
&lt;h4 id=&quot;integration&quot;&gt;&lt;a href=&quot;#integration&quot;&gt;Integration&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Testing solutions aren’t libraries that individual developers spin up and tear down when they have what they need. They need to be fully integrated into the enterprise software stack so they can enable easy access to data and trigger actions/alerts when tests pass or fail. And it is important to enable this in three ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;First, build primitives in the product that make it easy to integrate into other tools regardless of the enterprise stack—as each enterprise will have a unique collection of systems and tools.&lt;/li&gt;
&lt;li&gt;Second, invest in lightweight integrations that are more native for some of the most widely used infrastructure to make it even easier to get started in these cases.&lt;/li&gt;
&lt;li&gt;Third, provide implementation services that meet the needs of custom configurations. Each enterprise will have nuance in their needs around integrations, but it is critical that testing is fully integrated to deliver the most value.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&quot;lineage&quot;&gt;&lt;a href=&quot;#lineage&quot;&gt;Lineage&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;After tests have failed and the AI app has been debugged and re-productionalized, then various teams responsible for the AI application or related to AI governance need to be able to see what has happened and why. This information can’t be isolated in a developer notebook or trashed once the fix is made. There needs to be both persistence and provenance in this audit trail and a way to easily report on it to teams. This is important because teams need to know what was done to feel confident re-deploying the AI application.&lt;/p&gt;
&lt;p&gt;It’s also important for consistently managing reputational, regulatory, or operational risk. Teams don’t want to end up on the front page of the Journal for a chatbot going rogue—even less so finding themselves on the wrong end of a deca-million-dollar AI error.&lt;/p&gt;
&lt;h3 id=&quot;the-enterprise-platform-for-ai-testing&quot;&gt;&lt;a href=&quot;#the-enterprise-platform-for-ai-testing&quot;&gt;The enterprise platform for AI testing&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Enterprises need testing that is standard for AI applications in production. They need it to cut across all AI components to give visibility into what is actually causing the issue—not single usages of single models. They need multiple teammates with varying levels of interest in diving deep to be able to use or consume information off of the testing platform. They need a testing platform to be fully integrated with data sources, CI/CD pipelines, and alerts. And they need lineage on what has happened so they can audit this process and report on it to multiple constituencies.&lt;/p&gt;
&lt;p&gt;In short, testing is an enterprise problem that requires an enterprise solution. This is what we’ve built at Distributional, and I’d be happy to show you how it would work for you.&lt;/p&gt;
&lt;p&gt;If you are interested in learning more, here are a few ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Get access to our product&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;Reach out to our team&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://distributional.com/blog&quot;&gt;Read more on our blog&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded><category>archive</category></item><item><title>Why all ML and AI use cases need standardized testing</title><link>https://distributional.com/blog/why-all-ml-and-ai-use-cases-need-standardized-testing</link><guid isPermaLink="true">https://distributional.com/blog/why-all-ml-and-ai-use-cases-need-standardized-testing</guid><description>Part 2 of the enterprise AI testing series: the testing gaps in classical ML workflows, from process visibility to false-alarm suppression. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Thu, 24 Oct 2024 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;What do AI leaders &lt;em&gt;really&lt;/em&gt; think about AI today? We’ve had over 1,000 hours of conversations with AI leaders at more than 100 Fortune 500 enterprises. This is the second post in a three-part series summarizing lessons from these conversations, in which we’ll cover why leaders today believe that all AI/ML components—not just generative AI—need deeper testing. Read the first post on what makes AI so hard to test &lt;a href=&quot;https://distributional.com/blog/why-generative-ai-creates-unique-testing-challenges&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Before jumping in, it’s worth noting that I’m the CRO here at Distributional and the only one of our 25 team members without a technical degree. My take is therefore intended to be more accessible, less technical and, hopefully, a quick read. If you want to dive deeper, &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;let’s find time to talk&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;
&lt;h3 id=&quot;traditional-aiml-is-here-to-stay&quot;&gt;&lt;a href=&quot;#traditional-aiml-is-here-to-stay&quot;&gt;Traditional AI/ML is here to stay&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Although going from zero to one on Generative AI is the top priority for nearly every company we meet, traditional AI and machine learning is here to stay. Generative AI will certainly replace traditional AI/ML for certain tasks, but more often we see it being used to expand &lt;em&gt;the variety&lt;/em&gt; of tasks that AI handles or augment specific parts of existing AI/ML applications.&lt;/p&gt;
&lt;p&gt;Machine learning is particularly well suited to many of the tasks it handles today, especially in contexts where stability and explainability are important. In short, ML isn’t going anywhere. Time series, tabular data, and even some image and vision tasks will continue to be handled by more traditional machine learning or deep learning models independently.&lt;/p&gt;
&lt;p&gt;But these use cases are relatively mature, so a fair question is: do they need better testing methods? From our discussions with dozens of AI/ML team leaders, the answer is a resounding yes. Let’s dig into a few reasons why this is the case.&lt;/p&gt;
&lt;h3 id=&quot;ongoing-aiml-testing-challenges&quot;&gt;&lt;a href=&quot;#ongoing-aiml-testing-challenges&quot;&gt;Ongoing AI/ML testing challenges&lt;/a&gt;&lt;/h3&gt;
&lt;h4 id=&quot;visibility-into-process&quot;&gt;&lt;a href=&quot;#visibility-into-process&quot;&gt;Visibility into process&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Today, many AI/ML teams lack a clear way to understand application behavior on a global scale. “We have hundreds of models in production, but we don’t have a way to visualize how their behavior is shifting over time in a standard way, let alone in a single place,” said an AI product leader at an automotive company. “And this exposes us to outsized risk.”&lt;/p&gt;
&lt;p&gt;Even though AI/ML has been in production at scale for a decade at some companies, teams may still struggle to get visibility on behavior of these applications.&lt;/p&gt;
&lt;p&gt;To help with this, Distributional was designed to make it easy to unify how all behavioral metrics are collected, analyzed, tested and reported, and to do so in a customizable way so you are looking at a dashboard with only information, naming, and conventions that make sense to you.&lt;/p&gt;
&lt;h4 id=&quot;catching-real-issues&quot;&gt;&lt;a href=&quot;#catching-real-issues&quot;&gt;Catching real issues&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;All AI/ML applications have some degree of non-determinism and non-stationarity. Non-stationarity in particular is ratcheted up in production where both the application and the usage of the application could shift at the same time, making it hard to identify real issues and isolate where they are coming from.&lt;/p&gt;
&lt;p&gt;To catch these real issues with their AI applications, some teams monitor thresholds on summary statistics for their applications, but these can hide potential behavioral issues lurking in the broader distribution of usage data. Others only look at specific, recent windows of time, causing them to miss more subtle long-term shifts in behavior that may be worrisome.&lt;/p&gt;
&lt;p&gt;Distributional is designed to address these challenges with statistical tests on distributions of data and dynamic baselines that make it easy to explore multiple time windows for production AI systems. All of this is designed to make sure there isn’t anything hidden in your data—to fully test it to catch all of the lurking issues with an AI application in production.&lt;/p&gt;
&lt;h4 id=&quot;avoiding-false-alarms&quot;&gt;&lt;a href=&quot;#avoiding-false-alarms&quot;&gt;Avoiding false alarms&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Too many false alerts can make it impossible to identify which ones are real, and make it even harder to rally the team to address them. “We get so many alerts that we just don’t pay any attention to them anymore,” said an AI product manager at a large financial services company. “And it is hard to get the development team to take action on these issues if we don’t have a way to dynamically calibrate these tests or thresholds to fit the application, and show the development team the approach we took to do so.”&lt;/p&gt;
&lt;p&gt;AI/ML teams need new testing methods that help them find the signal within the noise. Distributional fills this gap with a workflow these teams can use to adapt their tests over time. They can automatically recalibrate tests over time at scale using reinforcement learning. And they can do deep root cause analysis directly within the same platform to analyze the specific data that is causing the issue.&lt;/p&gt;
&lt;h4 id=&quot;results-analysis&quot;&gt;&lt;a href=&quot;#results-analysis&quot;&gt;Results analysis&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;When AI is more mature and operating at a significant scale, this scale itself can become an issue. How do you do root cause analysis when you “stare at a wall of numbers,” as the technology leader at an automotive company recently told us?&lt;/p&gt;
&lt;p&gt;There are three things teams need to address this. First, teams need custom dashboards that are designed to help them understand the status of their unique AI applications. Then they need to know which test results are the cause of the issue. Finally, they need to be able to tag tests and filter data so they can explore various segments to determine what may actually be causing a particular issue. All of this helps to help these teams understand issues with their AI applications, triage these issues, and take appropriate action to resolve them before they create deeper problems in production.&lt;/p&gt;
&lt;h4 id=&quot;cross-team-workflow&quot;&gt;&lt;a href=&quot;#cross-team-workflow&quot;&gt;Cross-team workflow&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Large teams collaborating on ML/AI naturally run into complexity due to the scale of the operation. How do you get everyone on the same page? “We have one team that productionalizes models and houses all of the data, and another team who develops the ML models,” said an AI engineering leader at a large financial services company. “So when there is an issue, the team that productionalizes sends data to the team that developed the model, but then this development team lacks context to analyze the actual issue.”&lt;/p&gt;
&lt;p&gt;To solve this, you need to get multiple teams on the same page, working off of the same information and with the same workflow to calibrate tests or resolve issues. As a bonus, doing this well can build up better relationships between teams that can have a large halo effect on the broader organization. A better workflow starts with standardizing how tests are created, run, tracked and triaged, which requires a software platform–not a smattering of information across notebooks, docs, sheets and reports.&lt;/p&gt;
&lt;h4 id=&quot;reporting&quot;&gt;&lt;a href=&quot;#reporting&quot;&gt;Reporting&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Whether related to product efficacy, internal standards, or external regulations, governance is critical to company-wide adoption of AI. For GenAI, this entails designing a new process. But even for legacy AI/ML applications, there is often a meaningful gap in the information available to governance teams who need to analyze, resolve, and evolve AI test suites and results. An AI leader at a regulated company shared, “We send a point-in-time email with a set of daily eval results to our leadership team, but we need a way to do this systematically, continuously and with greater depth.” Another leader of continuous integration for traditional ML commented, “I need better ways to report out on lineage across tests and test results for internal and external compliance purposes.”&lt;/p&gt;
&lt;p&gt;AI testing isn’t just enabled with a better framework. It needs to include a full audit trail of issues identified and actions taken. And these tests need to connect to a dashboard with test results that can be analyzed and shared with various parties involved with AI governance. It is the power of putting all of these pieces together in a single software solution that can enable better AI decisions, reduce AI risk and allow for more sustainable AI governance.&lt;/p&gt;
&lt;h3 id=&quot;its-time-for-a-testing-upgrade&quot;&gt;&lt;a href=&quot;#its-time-for-a-testing-upgrade&quot;&gt;It’s time for a testing upgrade&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Generative AI is all the rage, but traditional AI and ML is here to stay. And these types of AI applications are often chained together in a single application pipeline, so testing solutions need to be able to cover all application types, not just generative AI.&lt;/p&gt;
&lt;p&gt;In summary, despite a decade of at-scale use, more traditional AI/ML applications still need a testing upgrade. Metrics are often computed and displayed by individual data scientists with no cross-cutting visibility in a unified dashboard for behavior. False negatives are often hidden in data you already have. A high volume of false positives can be hard to handle without a workflow that addresses them. It can be hard to parse results with current tooling, especially when operating at significant scale. Multiple teams are usually involved in the AI software lifecycle, and it can be hard for them to have a mutually productive workflow without a software platform in place that standardizes it. And being able to track a lineage and report on it is often just as important as resolving any given point-in-time issue.&lt;/p&gt;
&lt;p&gt;These challenges are hard, and deserve a better approach to AI testing. If you are interested in learning more, here are a few ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Get access to our product&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;Reach out to our team&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://distributional.com/blog&quot;&gt;Read more on our blog&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded><category>archive</category></item><item><title>Why generative AI creates unique testing challenges</title><link>https://distributional.com/blog/why-generative-ai-creates-unique-testing-challenges</link><guid isPermaLink="true">https://distributional.com/blog/why-generative-ai-creates-unique-testing-challenges</guid><description>Field notes from 1,000+ hours with Fortune 500 AI leaders on why GenAI breaks traditional testing: non-determinism, non-stationarity, and behavior definition. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Fri, 18 Oct 2024 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;What do AI leaders &lt;em&gt;really&lt;/em&gt; think about AI today? In a three-part series, we’re sharing the insights we’ve gleaned from over 1,000 hours of conversations with AI leaders at more than 100 Fortune 500 enterprises. This first post delineates the unique challenges companies face when they deploy generative AI, as compared with traditional software. The second covers why leaders today believe that all AI/ML components—not just generative AI—need deeper testing. And the final post focuses on why AI testing is an enterprise problem in need of an enterprise solution. (Read the second post &lt;a href=&quot;https://distributional.com/blog/why-all-ml-and-ai-use-cases-need-standardized-testing&quot;&gt;here&lt;/a&gt; &lt;em&gt;[NOTE: the original article&amp;#39;s link mistakenly pointed at this post itself; corrected to part 2.]&lt;/em&gt;).&lt;/p&gt;
&lt;p&gt;Before jumping in, it’s worth noting that I’m the CRO here at Distributional and the only one of our 25 team members without a technical degree. My take is therefore intended to be more accessible, less technical and, hopefully, a quick read. If you want to dive deeper, &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;let’s find time to talk&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;.&lt;/p&gt;
&lt;h3 id=&quot;generative-ai-is-on-the-rise&quot;&gt;&lt;a href=&quot;#generative-ai-is-on-the-rise&quot;&gt;Generative AI is on the rise&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Generative AI (GenAI) has real potential to transform the enterprise. Whether agents, copilots or apps designed for specific tasks, teams have started to deploy Large Language Models (LLMs) for both internal and external use cases.&lt;/p&gt;
&lt;p&gt;This opportunity is well established, but so are the challenges. Teams who have deployed GenAI struggle to detect and mitigate undesired behavior, resulting in hallucinations, incorrectness or an unreliable customer experience in production, among other issues. Other teams have a long backlog of these applications they want to deploy, but struggle with confidence in their behavior or capacity to satisfy governance needs—so these use cases are withering on the vine.&lt;/p&gt;
&lt;p&gt;No matter where you fall on this spectrum, a more complete approach to AI testing is one of the critical components to addressing these issues, much as traditional testing is for traditional software. But testing can mean different things to different people, so we spoke with over a hundred CIOs, CTOs, Chief AI Officers, VPs of AI, Directors of AI engineering and AI product managers to understand how they define AI testing. Digging deeper has yielded a few insights into how they think about this challenge—and some of the potential solutions.&lt;/p&gt;
&lt;h3 id=&quot;challenges-with-testing-ai&quot;&gt;&lt;a href=&quot;#challenges-with-testing-ai&quot;&gt;Challenges with testing AI&lt;/a&gt;&lt;/h3&gt;
&lt;h4 id=&quot;non-determinism&quot;&gt;&lt;a href=&quot;#non-determinism&quot;&gt;Non-determinism&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Non-determinism is when the same input can yield a multitude of possible outputs. For example, this might mean you prompt an LLM “What are your best recommendations for Florence?” and it might recommend a cooking class, a walking tour, or something else entirely each time you ask.
The level of variance within LLM responses is one of the characteristics that make them such powerful tools—it is ultimately a &lt;em&gt;desirable trait&lt;/em&gt;. But it can also make them hard to properly test and evaluate. The way Distributional tackles this problem is through testing on distributions of many inputs and outputs rather than single usages of an application. Any given usage of an LLM &lt;em&gt;should&lt;/em&gt; vary, but the distribution of usages should behave roughly the same. So tests need to be run on these distributions—something that Distributional was explicitly designed to do. Many companies still face this challenge today. Prior to trying Distributional, a CTO at a data services company told us, “Non-deterministic apps need to be continuously checked to ensure they’re at a steady state based on our use case, and we lack solutions to do this well today.”&lt;/p&gt;
&lt;h4 id=&quot;non-stationarity&quot;&gt;&lt;a href=&quot;#non-stationarity&quot;&gt;Non-stationarity&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Most AI, GenAI or traditional ML applications include non-stationary components, but this is particularly pervasive for LLMs. The entire application could be managed by a third-party vendor. The LLM powering an application could be a third-party-managed API. There could be evolving datasets supporting RAG applications or driving financial risk scoring. The underlying infrastructure running these applications may shift underneath you as teams update their models. These upstream shifts can cause shifts in the behavior of the application that relies on it, even if nothing was changed in the application itself. All of these are instances of non-stationarity that require all data of these AI/ML applications—including for their upstream components—to be continuously and adaptively tested over time.&lt;/p&gt;
&lt;h4 id=&quot;defining-behavior&quot;&gt;&lt;a href=&quot;#defining-behavior&quot;&gt;Defining behavior&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;“Many GenAI tasks are subjective, so how do you measure its behavior? Are judge metrics reliable? Do you automatically compute these off of raw text?” asked an ML engineer at a large consumer technology company recently. Most conversations we’ve had with our customers include some variation of these questions. GenAI is so early that most teams are still parsing which metrics are the best representation of model performance, let alone behavior. To streamline this process, teams use Distributional to automatically derive many performance and behavioral properties off of your text data, adding a slew of information about model behavior that you can test alongside any custom evals you’ve developed. This creates a rich representation of behavior from potentially limited application data. Our platform then automatically recommends tests and adaptively calibrates these tests to fit your unique AI application ensuring this behavior doesn’t deviate over time.&lt;/p&gt;
&lt;h4 id=&quot;pipelines&quot;&gt;&lt;a href=&quot;#pipelines&quot;&gt;Pipelines&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;AI applications rarely exist in isolation. They depend on upstream components that produce features, third party packages that may need to be upgraded, third party data sources that shift over time, or third party APIs, such as hosted LLM endpoints. In most cases, there are typically multiple non-stationary components in any given AI or ML pipeline. And if non-stationarity or non-determinism exists anywhere, it propagates through these systems. AI leaders are aware of this issue. One AI leader at a consumer electronics company told us, “I run evals that give us a good sense of how our core app performs today, but we rely on upstream dependencies that shift over time and need a way to standardize how tests are run across each of these components as well.” The entire computational graph representing an application needs to be tested in unison to effectively root cause the origin of any potential issue or behavioral change.&lt;/p&gt;
&lt;h4 id=&quot;democratization&quot;&gt;&lt;a href=&quot;#democratization&quot;&gt;Democratization&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Growing use of pre-trained models—and LLMs in particular—makes it easier than ever for anyone to develop their own AI applications. Yet a lack of confidence is holding companies back from reaping the full benefits. One CTO of a large insurance company told us, “This is the first time in my life where we’re getting pull from business lines to use a specific technology. But unless I standardize testing and other operations around LLMs, we’re not in a position to enable any of these use cases.”&lt;/p&gt;
&lt;p&gt;Leaders want to empower teams to unlock LLM use cases for their business lines, but also know that doing so without proper checks risks reputational, operational or regulatory harm. There is a huge opportunity to democratize AI development, but it must come with standardized enterprise tooling that can provide repeatability, consistency and visibility to the process. The goal is to enable flexibility in terms of what is being developed, but then to verify it fits standards the company sets through testing.&lt;/p&gt;
&lt;h4 id=&quot;opportunity-cost&quot;&gt;&lt;a href=&quot;#opportunity-cost&quot;&gt;Opportunity cost&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;The challenges of building reliable AI today comes with a massive opportunity cost for companies. A VP of AI at a financial technology company told me, “With traditional ML, development took the vast majority of my team’s time. Now with LLMs, testing and evaluation takes 5-10x as much time as development.” Similarly, another CPO of a large enterprise technology company shared, “I know we need better testing for our GenAI, but we are prioritizing building new revenue-generating AI features instead.”&lt;/p&gt;
&lt;p&gt;In all cases, teams need a workflow to automate how testing and validation is done on their applications so this step takes less valuable team time. This is why at Distributional we’ve prioritized automating the process from data collection to augmentation to testing to adaptive recalibration in our product to ensure AI teams can quickly reap the benefits of testing to get more confidence in their AI applications.&lt;/p&gt;
&lt;h3 id=&quot;a-challenge-that-distributional-solves&quot;&gt;&lt;a href=&quot;#a-challenge-that-distributional-solves&quot;&gt;A challenge that Distributional solves&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;In summary, because generative AI is such a large opportunity for most companies, it also carries significant risk. It’s prone to non-determinism that doesn’t allow for traditional software testing. It often includes many non-stationary components with varying levels of team control. It is also so new that there isn’t a standard set of metrics teams can align on. LLMs are often chained together or embedded in pipelines with other ML components, so any non-determinism or non-stationarity propagates through the entire system. With pre-trained LLMs and LLM API endpoints, it’s easier than ever for a business line to develop an AI app, but also easier than ever for them to ship without a clear concept of AI app behavior. And proper testing is both time-intensive and orthogonal to an AI team’s daily responsibilities. Distributional is building the solution to all of these challenges.&lt;/p&gt;
&lt;p&gt;If you are interested in learning more, here are a few ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://distributional.com/sign-up/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Get access to our product&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;Reach out to our team&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://distributional.com/blog&quot;&gt;Read more on our blog&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded><category>archive</category></item><item><title>We raised $11M for better AI testing</title><link>https://distributional.com/blog/we-raised-11m-for-better-ai-testing</link><guid isPermaLink="true">https://distributional.com/blog/we-raised-11m-for-better-ai-testing</guid><description>Scott Clark&apos;s 2023 founding essay: a decade of AI testing problems through Yelp, SigOpt, and Intel, and the thesis behind Distributional. From the Distributional archive. The product described has been sunset: read the pivot post, then the Talaria Scientific manifesto.</description><pubDate>Thu, 14 Dec 2023 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;&lt;p&gt;Archive note: From the Distributional archive: this post is preserved with its original byline and date, and the product it describes has been sunset. Distributional is now Talaria Scientific. Read &lt;a href=&quot;https://distributional.com/blog/distributional-is-now-talaria&quot;&gt;the pivot post&lt;/a&gt; first, then &lt;a href=&quot;https://talariasci.com/blog/why-im-building-talaria&quot;&gt;the Talaria manifesto&lt;/a&gt;.&lt;/p&gt;&lt;/blockquote&gt;&lt;h3 id=&quot;summary&quot;&gt;&lt;a href=&quot;#summary&quot;&gt;Summary&lt;/a&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;As the capacity of AI across enterprise tasks grows, so does its potential risk to these businesses and their customers. Every day there is a new report of AI bias, instability, failure, error or other issues.&lt;/li&gt;
&lt;li&gt;This is a problem with massive scale. Marc Andreessen has called &lt;a href=&quot;https://www.youtube.com/watch?v=0wIUK0nsyUg&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;AI correctness and security&lt;/a&gt; trillion-dollar software problems.&lt;/li&gt;
&lt;li&gt;Distributional is building the modern enterprise AI testing and evaluation platform designed to enable our customers to identify, understand and address AI risk before AI-enabled products are deployed.&lt;/li&gt;
&lt;li&gt;Distributional has an 11-person founding team led by Scott Clark, co-founder and CEO of SigOpt, acquired by Intel in 2020, as well as a team of AI, platform and research engineers from Bloomberg, Google, Intel, Meta, SigOpt, Slack, Stripe, Uber and Yelp.&lt;/li&gt;
&lt;li&gt;To fuel our product vision, we are announcing a $11M Seed round led by Andresseen Horowitz with participation from Operator Stack, Point72 Ventures, SV Angel, Two Sigma and Willowtree Investments.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;introducing-distributional&quot;&gt;&lt;a href=&quot;#introducing-distributional&quot;&gt;Introducing Distributional&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;“How do you test these models today?” I asked the head of the AI platform engineering team at a company that relies on thousands of models in production as part of its core business.&lt;/p&gt;
&lt;p&gt;“We have over 500 engineers and analysts who are responsible for deep testing and retesting of every model on a daily basis. If these models shift, they are responsible for finding, evaluating and fixing these issues.”&lt;/p&gt;
&lt;p&gt;“Do they have any standardized tools to do this work systematically? Do you aspire to this?”&lt;/p&gt;
&lt;p&gt;“No, they each choose their own approach. And, yes, we would like to automate testing but haven’t found the right approach yet.”&lt;/p&gt;
&lt;p&gt;In recent months, I have had what feels like the same conversation with dozens of AI leaders in finance, technology, energy, semiconductors, pharmaceuticals, consulting, software and manufacturing. AI – whether traditional machine learning, deep learning, generative AI or the large language models (LLMs) dominating the generative space – is complex, often unpredictable, and constantly changing. Whether from hallucinations, instability, inaccuracy, integration or dozens of other potential challenges, these teams struggle to identify, understand and address AI risk with depth or at scale.&lt;/p&gt;
&lt;p&gt;I am often astonished by the differences between traditional software engineering and AI-enabled software development. Testing is standard for traditional software. Teams try to maximize the coverage of their tests and root out “flaky” tests that spuriously fail as they bridge the gap between development and deployment. Engineering teams run unit, regression and integration tests in scalable CI/CD pipelines before putting code in production.&lt;/p&gt;
&lt;p&gt;But when AI is added, introducing more math and randomness, the complexity of testing these systems explodes on a variety of dimensions at once. Standard tools no longer work for your purpose. AI models often are given a pass because they are “too complex” and unpredictable. This causes too much uncertainty so coverage is no longer enough. Proper AI testing requires depth, which is a hard problem to solve.&lt;/p&gt;
&lt;p&gt;This is in part why AI has been described as &lt;a href=&quot;https://research.google/pubs/pub43146/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;the high interest credit card of technical debt&lt;/a&gt;, A huge part of this debt is insufficient testing. Most teams choose to assume model behavior risk, and accept that models will have issues. Some may try ad-hoc manual testing to find these issues, which is often resource intensive, disorganized, and inherently incomplete. Others may try to passively catch these issues with monitoring tools after AI is in production. In many cases, teams choose to avoid AI even when it could be useful for their applications. In all cases, these teams know there is significant risk around their AI-enabled applications and that they need more robust testing to understand and address it. And, increasingly, these teams may also be required to do this through &lt;a href=&quot;https://www.nytimes.com/2023/09/12/technology/white-house-ai-tech-pledge.html&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;shareholder pressure&lt;/a&gt;, &lt;a href=&quot;https://www.whitehouse.gov/briefing-room/statements-releases/2023/10/30/fact-sheet-president-biden-issues-executive-order-on-safe-secure-and-trustworthy-artificial-intelligence/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;government regulation&lt;/a&gt; or &lt;a href=&quot;https://www.nist.gov/artificial-intelligence/executive-order-safe-secure-and-trustworthy-artificial-intelligence&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;industry standards&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;We founded Distributional to solve this problem. Our mission is to empower our customers to actively make their AI-based products more safe, reliable, and secure, before they deploy them. We aim to catch harm before their customers or users do.&lt;/p&gt;
&lt;p&gt;To pursue this mission, we raised an $11M seed led by Andreessen Horowitz with Martin Casado joining the board and with participation from Operator Stack, Point72 Ventures, SV Angel, Two Sigma, Willowtree Investments, and more than 40 other AI leaders in industry and academia as angel investors. In a &lt;a href=&quot;https://www.youtube.com/watch?v=0wIUK0nsyUg&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;recent interview&lt;/a&gt; with Martin, Marc Andreessen said, “to make AI generally useful in a way that is guaranteed to be correct or secure – these are two of the biggest opportunities I’ve ever seen in my career.” Armed with their deep support and expertise, our &lt;a href=&quot;https://www.linkedin.com/company/dbnlai/people/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;founding team of 11&lt;/a&gt; is poised to realize this opportunity.&lt;/p&gt;
&lt;h3 id=&quot;a-decade-of-ai-testing-problems&quot;&gt;&lt;a href=&quot;#a-decade-of-ai-testing-problems&quot;&gt;A Decade of AI Testing Problems&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Our partners, customers, and investors give our team a broad perspective on this problem. But what makes our perspective unique is that we combine their insights with our direct experience attempting to solve versions of this problem for nearly a decade.&lt;/p&gt;
&lt;h4 id=&quot;2014-ai-evaluation&quot;&gt;&lt;a href=&quot;#2014-ai-evaluation&quot;&gt;2014: AI Evaluation&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;We first saw this problem in our own software at &lt;a href=&quot;https://sigopt.com/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;SigOpt&lt;/a&gt;, the AI startup I previously founded, when building our optimization and experimentation platform for enterprise scale in 2014. We had developed cutting edge ways to efficiently optimize complex systems, but were constantly exploring new techniques to improve this solution. To feel confident deploying new algorithmic solutions, we needed to rigorously test them and have confidence in their robustness.&lt;/p&gt;
&lt;p&gt;We considered A/B testing, but we couldn’t run these tests in production due to the risk of real customer harm. Not to mention, this approach was antithetical to our value proposition of extremely efficient optimization. We also considered standard frameworks for benchmarking optimization methods, but couldn’t find one designed to robustly compare results from stochastic methods. With no available solution, our team instead built an evaluation framework and &lt;a href=&quot;https://proceedings.mlr.press/v64/dewancker_strategy_2016.html&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;published&lt;/a&gt; it at the ICML workshop on optimization in 2016.&lt;/p&gt;
&lt;p&gt;Being able to confidently claim we had the best, and most tested, optimization framework became one of our strongest competitive advantages in the years to come. More importantly, this evaluation process exposed valuable insights on product performance. It was often shocking what methods looked good in a paper, but did not perform well when exposed to rigorous testing. By continuously testing we were able to cut out poor performing techniques before they ever made it to our users. Although we were proud of our invention, even our team believed it would have been great to use standardized tooling here instead of needing to build it ourselves from scratch.&lt;/p&gt;
&lt;h4 id=&quot;2016-2020-ai-robustness&quot;&gt;&lt;a href=&quot;#2016-2020-ai-robustness&quot;&gt;2016-2020: AI Robustness&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;After we established SigOpt as a reliable, sample-efficient product for optimizing black box systems, our product was increasingly used by sophisticated companies deploying AI as a core component of their product or revenue strategy. These teams had high upside for boosting performance of their models, but also significant downside if they didn’t perform as expected. So they often valued robustness as much or more than performance.&lt;/p&gt;
&lt;p&gt;For example, if one of our clients were to utilize a brittle model to make important business decisions, subtle shifts in inputs may lead to widely varying outputs and suboptimal outcomes. As SigOpt made these models better and more powerful, the need for robustness – and the tradeoff between robustness and maximum potential performance – became more important.&lt;/p&gt;
&lt;p&gt;It is often better to have a solution at 90% of perfect all the time than a solution that wildly oscillates between 99% and 10%. This is a very difficult problem in high dimensions of input and output where traditional perturbation analysis is prohibitively expensive.&lt;/p&gt;
&lt;p&gt;Once you find an optimal model, how can you evaluate whether it is brittle? And how do you make sure this understanding of optimal performance and relative brittleness doesn’t shift over time?&lt;/p&gt;
&lt;p&gt;As we saw the rise of this use case across our user base, we designed a purpose built solution to this problem called &lt;a href=&quot;https://sigopt.com/blog/highlight-constraint-active-search/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Constraint Active Search&lt;/a&gt; and &lt;a href=&quot;http://proceedings.mlr.press/v139/malkomes21a/malkomes21a.pdf&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;published&lt;/a&gt; it at ICML 2021. This algorithmic technique allowed these teams to set constraints on a variety of metrics and run experiments that would actively probe and produce a variety of performant models that satisfied these constraints. Users loved this feature because it allowed them to effectively and efficiently optimize their model reliably against different permutations of input parameters in ways they never could before. In turn, they built more intuition on model robustness and had more confidence that the model they deployed wouldn’t significantly degrade with shifts in input distributions.&lt;/p&gt;
&lt;h4 id=&quot;2022-continuous-ai-testing-at-scale&quot;&gt;&lt;a href=&quot;#2022-continuous-ai-testing-at-scale&quot;&gt;2022: Continuous AI Testing at Scale&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;In October 2020, Intel acquired SigOpt. At Intel, I had the privilege of leading the AI and HPC software teams in the Supercomputing division that was bringing Intel’s next generation of GPUs and HPC-oriented CPUs to market. In this role, I managed over one hundred engineers with the purpose of running, evaluating, debugging and evolving AI and HPC workloads for each new processor we were bringing to market. Given the sophistication of our customers, most of this work involved complex AI and physical modeling. This process translated to our teams orchestrating up to thousands of AI test workloads daily.&lt;/p&gt;
&lt;p&gt;As we built out the full software stack for this task, there were robust frameworks in place for traditional software testing, but nothing similar for AI.  As a result, our team was forced to spend most of its time and energy manually designing, instrumenting, executing and analyzing tests for AI workloads. We explored options for supporting this workflow with software, but couldn’t find a robust enough solution or a reliable testing framework. Although we had ambitions for continuous testing, this simply wasn’t attainable without automation in place. One member of the executive team called AI testing a “million dollar per day problem for companies operating at this size and scale.” This was a huge problem, but there were no good off-the-shelf solutions internally or externally to address it.&lt;/p&gt;
&lt;h3 id=&quot;better-testing-greater-ai-impact&quot;&gt;&lt;a href=&quot;#better-testing-greater-ai-impact&quot;&gt;Better Testing, Greater AI Impact&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Through conversations with AI product leaders in finance, energy and tech, I have come to realize that these are quite common issues. These leaders agree that software requires testing, but traditional testing methods and frameworks were built around assumptions that do not hold for applications built on AI.&lt;/p&gt;
&lt;p&gt;Engineers are often forced to test these models by partially fitting them into legacy testing frameworks (often only testing metric thresholds and summary statistics), applying qualitative analysis as they build models (using visualizations or hand-constructed examples in notebooks to gain intuition and confidence), or shifting their problem to their users and customers by letting them test it live via online monitoring. As a consequence, AI is incompletely and non-continuously tested today. This exposes the business to significant risk, high opportunity cost, or both.&lt;/p&gt;
&lt;p&gt;When I ask CIOs how they know that their models haven’t introduced bias or gone off the rails they often say that their only recourse is to constantly monitor metrics and feedback, passively waiting for something to go wrong. While monitoring is an important part of any software product, having the confidence of rigorous, active testing allows product teams to deploy without fear that something catastrophic is right around the corner and only discoverable by your users, after the fact.&lt;/p&gt;
&lt;p&gt;As AI becomes more powerful and prevalent, it becomes increasingly important to make sure it is tested and performing as expected. We hope to usher in a virtuous cycle for our customers. With better testing, teams will have more confidence deploying AI in their applications. As they deploy more AI, they will see its impact grow exponentially. And as they see this impact scale, they will apply it to more complex and meaningful problems, which in turn will need even more testing to ensure it is safe, reliable, and secure.&lt;/p&gt;
&lt;h3 id=&quot;whats-next&quot;&gt;&lt;a href=&quot;#whats-next&quot;&gt;What’s Next?&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;We aspire to help our customers realize this future and would love your help along the way. We are collaborating with more than a dozen co-design partners to build the modern enterprise platform for AI testing, but are always interested in expanding the scope of our collaboration.&lt;/p&gt;
&lt;p&gt;Here are ways to get involved:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://distributional.com/sign-up&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Sign up&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; for early access to our private beta&lt;/li&gt;
&lt;li&gt;Let us know your &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;interest&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; in joining the team&lt;/li&gt;
&lt;li&gt;Read &lt;a href=&quot;https://a16z.com/announcement/investing-in-distributional/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;this post&lt;/a&gt; on the market opportunity from Martin Casado at a16z&lt;/li&gt;
&lt;li&gt;Read &lt;a href=&quot;https://p72.vc/perspectives/our-investment-in-distributional/&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;this post&lt;/a&gt; on the product problem from Noah Carr at Point72 Ventures&lt;/li&gt;
&lt;li&gt;Follow us on &lt;a href=&quot;https://www.linkedin.com/company/dbnlai&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;https://twitter.com/dbnlai&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;X/Twitter&lt;/a&gt; and &lt;a href=&quot;https://www.youtube.com/@distributionalai&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Youtube&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Reach out to &lt;a href=&quot;mailto:contact@distributional.com&quot;&gt;share&lt;/a&gt; &lt;em&gt;[NOTE: Distributional has been sunset; Distributional is now &lt;a href=&quot;https://talariasci.com&quot; rel=&quot;noopener noreferrer&quot; target=&quot;_blank&quot;&gt;Talaria Scientific&lt;/a&gt;. This link is preserved for the historical record.]&lt;/em&gt; your perspective&lt;/li&gt;
&lt;/ul&gt;</content:encoded><category>archive</category></item></channel></rss>