What we still need to verify : 2 points in this profile are not yet confirmed against vendor documentation.
- Feature boundary between the open source component and the commercial platform: confirm
- Which safety and security evaluators ship built in: verify against vendor docs
Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.
What it does
Arize instruments AI applications and collects what they do. Instrumentation is built on OpenTelemetry with a semantic convention for LLM workloads, so a trace captures a request end to end: the prompt as sent including retrieved context, the model and parameters, tool calls with their arguments and results, token counts, latency, and the final response. For agentic applications the trace is a tree, and the nested spans show which step produced a bad answer.
On top of the traces sits an evaluation layer. Evaluators run over recorded spans and score them, typically using a language model as judge against a rubric, covering answer relevance, retrieval quality, hallucination, and whether a response contains content it should not. That mechanism is what gives the platform a security role: evaluate inputs for injection attempts and outputs for leaked sensitive data, then alert on rates rather than individual events. Datasets and experiments let you pin inputs, replay them against a changed prompt or model, and compare scores. The open source component provides tracing and evaluation you can run locally or self host, with the commercial platform adding managed storage, monitoring and collaboration.
Where it fits
Development and production both. Engineers use the local tracing view while building, and the same instrumentation feeds a hosted or self hosted backend once the application ships. Ownership usually sits with the AI engineering team, not security: security gets value here by defining evaluators and consuming the metrics, not by operating the tool. You need to instrument the application, and you need to decide what a bad response looks like before an evaluator can find one.
Strengths
- OpenTelemetry based instrumentation avoids a proprietary agent and slots into existing observability pipelines.
- Trace trees make agent and retrieval failures diagnosable at the step that caused them.
- Dataset and experiment workflow turns prompt changes into measured comparisons rather than vibes.
- A usable open source path exists for teams that cannot send prompt data to a vendor.
Limitations
- This is an observability platform first. Security detection is something you configure with evaluators, not a hardened control, and it does not block anything in line by itself.
- LLM as judge evaluation costs money per span and carries its own error rate, so evaluating full production volume is usually impractical and sampling hides rare events.
- Traces contain prompts and retrieved context, which means they contain whatever sensitive data your application handles. Retention and access control become a real obligation.
Who it suits
Teams operating LLM features at enough volume that behavior must be measured rather than assumed, with engineering capacity to instrument and to write evaluators. Not a substitute for a runtime guardrail if your requirement is blocking bad input or output rather than observing it.
Used Arize AI? Recommend it under your own name and title.
Recommend this tool