Creuto is now an OpenAI Select Partner Read More

Software Architecture & Technical

Dynatrace Arize: $915M buys a verdict, not more traces

The Dynatrace Arize deal closed at $915M. Why agent failures look healthy in a trace, where Arize meets OpenTelemetry, and what to do whatever you run.

Dynatrace Arize: $915M buys a verdict, not more traces

Arize's own documentation describes the failure its customers pay for in one sentence: "Your agent returns a 200, the trace looks fine, and the user gets a bad answer." The Dynatrace Arize deal closed on 1 October 2026 at $915 million, and that sentence is what the money bought — not more telemetry, but a verdict on whether a decision was right. If you run agents in production, the useful part of this news is not the price. It is the admission that spans, error rates and latency percentiles were never going to catch the failure mode you actually have.

Here are the terms, from the primary source rather than the trade write-ups:

  • $915 million total, consisting of approximately $815 million in cash plus replacement equity awards for Arize employees joining Dynatrace, funded from cash on hand and/or the existing credit facility (Dynatrace investor release, 13 August 2026).
  • Both Arize founders, Jason Lopatecki and Aparna Dhinakaran, join at closing; Lopatecki continues to lead the Arize team and reports to Dynatrace CEO Rick McConnell.
  • Dynatrace expects the transaction to be roughly 200 basis points accretive to ARR growth and 175 basis points dilutive to non-GAAP operating margin for fiscal 2027.
  • The deal was announced on 13 August 2026 and completed on 1 October 2026. Arize continues to support both Phoenix, its open-source platform, and AX, its enterprise platform, with capabilities folded into Dynatrace "over time".

Why the Dynatrace Arize deal is an evaluation story, not a telemetry one

Traditional application performance monitoring answers two questions well: was it slow, and did it throw. Both are properties of the call, and both are cheap to compute from a span. An agent breaks a third way. It completes, returns 200, stays inside its latency budget, and gives an answer that is wrong — or right for a reason that will not hold next week.

Arize's documentation names the distinction more plainly than the press release does: "A trace tells you what your AI agent did. An eval tells you whether what it did was good." The same page lists the failure modes it is built for — confident falsehoods, plausible reasoning that lands on the wrong answer, the wrong context retrieved, the wrong tool called, a prompt change that quietly degrades quality for users whose traces nobody reads. None of them raises an error.

That is a different discipline from the one observability vendors have been selling, because it needs something telemetry does not contain: a definition of correct. Spans record what happened. Deciding whether what happened was good requires a judge, a rubric and ground truth — which is why scoring the trajectory rather than the final answer turns out to be the harder half of agent work, and why the judge itself needs correcting before you trust it. We have written separately about the biases an LLM judge brings to its own scoring.

QuestionWhat answers itWhat it costs you
Was it slow?Span duration, latency percentilesAlready instrumented
Did it throw?Status codes, error ratesAlready instrumented
Was the decision right?An evaluator you wrote, run against criteria you definedHuman review time, judge tokens, and the argument about what "right" means

Where Arize overlaps OpenTelemetry, and where it does not

Arize did not build a parallel tracing stack. OpenInference, its open specification, describes itself as "complementary to OpenTelemetry" and works with any OpenTelemetry-compatible backend. It adds AI-specific semantic conventions on top of ordinary spans: a required openinference.span.kind attribute whose values include LLM, RETRIEVER, TOOL, AGENT, GUARDRAIL and EVALUATOR. Phoenix, the open-source product, is built on OpenTelemetry and describes itself as vendor, language and framework agnostic. If you have already done an OpenTelemetry migration, the transport and the collector stay where they are.

The overlap now extends further than it did. OpenTelemetry's GenAI semantic conventions, which moved into their own repository, already register gen_ai.evaluation.name, gen_ai.evaluation.score.value, gen_ai.evaluation.score.label and gen_ai.evaluation.explanation — all at Development status. So the standard has a place to put a verdict. It does not supply the verdict, the rubric, or anyone to argue with about them.

Where the models genuinely diverge is time. OpenInference's annotation spec states the constraint directly: ended spans are immutable, so feedback that arrives after a span closes must travel on a new carrier span holding exactly one OpenTelemetry span link back to the target. Telemetry is written as the work happens. A judgement usually arrives minutes or hours later, from a human reviewer or a batch judge. An evaluation platform is, in large part, machinery for attaching late verdicts to finished work — which is not something APM was ever asked to do.

The volume problem is a sampling problem in disguise

The quote travelling with this deal is Dhinakaran's, speaking to The New Stack on 1 October 2026: "No human wants to go look at billions of traces. Nobody's going to go do that." True, and it is usually read as an argument for agents that read telemetry for you. The sharper consequence is upstream of that.

OpenTelemetry's sampling guidance tells you to consider sampling when "most of your trace data represents healthy traffic with little variation" and when you have "common criteria, like errors or high latency, that usually means something is wrong". For high-volume systems it notes that 1% or lower can accurately represent the other 99%. Every one of those assumptions is sound for request telemetry and wrong for agents. The healthy-looking traces are exactly where the bad answers hide, and head sampling cannot inspect a trace it has already decided to drop. Tail sampling helps, but it keys on errors, latency and attributes — and "the answer was wrong" is not an attribute at sampling time unless something already judged it.

Arize's own guidance concedes the economics rather than pretending they vanish. Its documentation recommends evaluating 100% of traces only for low-volume or critical applications, 10–50% for high-volume ones, and 1–5% where representative sampling is enough, starting at 10–20% until you have validated the evaluator. The interesting mechanism is not the percentage but the filter: multi-span queries admit traces matching a structural pattern — retrieval happened and then an LLM answered, say — before sampling applies. Choosing which traces are eligible to be judged beats choosing a smaller random fraction of everything.

The strongest case against buying any of this

It goes like this. You already run OpenTelemetry. The conventions for recording an evaluation score on a span are in the standard. You can write a judge in an afternoon, emit its score as a span attribute, tail-sample on that attribute and chart it in whatever backend you already pay for. Evaluation is a feature, not a category, and a $915 million price tag is consolidation economics rather than engineering necessity.

Most of that is correct, and it is the right instinct if your agent surface is small. What it underestimates is which part is expensive. The plumbing is a weekend. The durable cost is the rubric: deciding what a good answer looks like for your product, finding the failure modes by reading traces by hand, encoding each one as an evaluator, then checking that the judge agrees with your humans often enough to be trusted on the traces nobody reads. Arize's own documentation is explicit that human review and automated evals do different jobs — review finds what is broken and defines "good"; evals assess quality at scale — and that you need both. No acquisition changes that, and no platform choice is permanent if the criteria live in your repository rather than in a vendor's hub.

What to do now, whatever platform you run

The deal will be old news in a fortnight. This part is not:

  1. Read fifty real traces by hand before you buy anything. You cannot encode failure modes you have not seen, and the list is always shorter and stranger than the generic eval catalogue suggests.
  2. Write the criteria down as versioned artefacts in your own repository, not as prompts typed into a vendor UI. They are the asset. The runner is replaceable.
  3. Measure your judge against human labels before you trust its scores on unreviewed traffic, and re-check it after every model change.
  4. Sample deliberately, not randomly. Filter to the span patterns that carry risk — tool calls that write, retrieval that feeds a customer-facing answer — and spend your evaluation budget there.
  5. Budget the judge. Evaluation is inference, so it carries its own token bill; the same traces, caps and cost controls that govern your agents have to govern the evaluators too.

What customers of either product should watch

Neither company has published an integration roadmap, so the honest advice is about what to verify rather than what to predict. Dynatrace's completion release says Arize will continue to support both Phoenix and AX, and that Arize capabilities will be integrated into Dynatrace over time; Dhinakaran told The New Stack that Arize remains available standalone, so customers will not need to be Dynatrace users. Both statements are worth holding on to in writing at your next renewal.

Three things to check, in order. Whether your instrumentation stays exportable — OpenInference spans are OpenTelemetry spans, and that is your exit. Whether packaging changes where the two products meet, since evaluation volume and observability ingest have historically been priced on different units. And whether the open-source path keeps parity with the commercial one, which is the usual casualty of this kind of deal and the easiest thing to test: instrument once, point the exporter at both, and see what arrives.

If you are deciding this quarter, the sequence matters more than the vendor. Define what a good answer is for your product, encode it, check the judge, then pick a platform to run it on — because the platform you pick is the part you can change later. Our DevOps and cloud engineering work starts in that order for the same reason.

Frequently asked questions

Dynatrace acquired Arize for $915 million, consisting of approximately $815 million in cash plus replacement equity awards for Arize employees joining Dynatrace. The deal was announced on 13 August 2026 and completed on 1 October 2026, according to Dynatrace's investor relations releases.

Arize is an AI observability and evaluation platform for LLM and agent applications. It traces agent behaviour, runs evaluations that score outputs against criteria you define, and supports two products: Phoenix, its open-source platform, and AX, its enterprise platform.

LLM observability answers a question APM does not ask. APM reports latency and errors, which are properties of a call. An agent can return 200 inside its latency budget and still give a wrong answer, so judging quality needs an evaluator and a rubric, not telemetry.

Not necessarily. OpenInference spans are OpenTelemetry spans and work with any OpenTelemetry-compatible backend, and OpenTelemetry's GenAI conventions already register evaluation score attributes. A dedicated platform buys the evaluation workflow, not the transport, so the decision is about workflow.

Dynatrace's completion release says Arize will continue supporting both Phoenix and AX, with Arize capabilities integrated into Dynatrace over time. Neither company has published an integration roadmap, so customers should confirm export paths, packaging and open-source parity at renewal.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

2 Oct 2026

·

9 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved