New Relic Introduces AI Evaluation to Close the Loop Across the Entire Developer-to-Production Lifecycle with Transaction-level Business Impact
AI Evaluation is an advanced evaluation framework that brings SREs, developers and engineers the larger context for all
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.
![]()
NEW RELIC NOW—New Relic, the intelligent observability platform, today announced a framework for advanced AI evaluation. New Relic AI Evaluation is a new capability within New Relic AI Observability. Unlike point solutions that evaluate isolated single-LLM calls, it delivers transaction-level business impact insights across the full developer-to-production lifecycle, tracking response quality and behavior while providing automated, real-time insights into AI guardrail performance.
Delivering end-to-end trust: Full-stack visibility built for the AI era
As generative AI becomes deeply embedded across enterprise workflows, software is shifting from deterministic code to probabilistic systems that fail in various ways, from subtle hallucinations to unexpected performance shifts when frontier vendors update LLMs. This creates a new observability requirement as traditional application signals no longer tell the whole story. Teams now need to see which model responded, which prompt was used, what tools an agent called, how much the interaction cost, and whether the result was useful. Observing AI in a silo or through best-of-breed point solutions creates dangerous blind spots around upstream and downstream operational impacts while widening the visibility gap between AI developers and production engineering. A single, unified intelligent observability platform built for the AI era solves this by uniting full-stack operational telemetry with real-time AI evaluations—delivering the essential trust layer and end-to-end visibility leaders need to control costs, mitigate security risks, and safely scale enterprise AI.
Built natively into the unified New Relic platform, AI Evaluation goes beyond isolated prompt checks to analyze AI performance across the entire application lifecycle—down to the underlying transaction. By uniting full context rather than focusing on a single metric, it embeds quality and security guardrails directly into existing application performance monitoring workflows.
Replacing unscalable manual reviews with an asynchronous “LLM-as-a-judge” service, AI Evaluation automatically scans live telemetry to score vulnerabilities like hallucinations, prompt injections, and data leaks. It attaches these probabilistic quality scores as attributes directly to deterministic distributed traces, allowing teams to isolate the exact root cause of a failure, whether in the prompt, vector database, or backend infrastructure, within a single view. By linking qualitative response scores directly to underlying compute consumption, teams can easily determine if expensive models deliver enough semantic value over faster, lower-cost alternatives.
“Generative AI has redefined application health, making evaluation of response quality and model efficiency just as critical as uptime,” said New Relic Chief Product Officer Brian Emerson. “Unlike point solutions that evaluate isolated LLM calls, we look holistically across the entire application transaction to show exactly what happens whenever AI is involved, giving teams full visibility into technical health, business impact, and performance—all within the platform tools that SREs, platform engineers, and developers already use.”
Key features and benefits of the new capabilities include:
Real-Time AI Evaluation and Guardrails:
- Direct Trace Integration for Rapid Troubleshooting: AI Evaluation attaches probabilistic quality scores to core deterministic distributed traces. Instead of chasing bad answers through siloed logs, engineers can view the prompt payload, judge reasoning, and cross-stack application behavior in a single screen to instantly isolate and resolve issues.
- Automated Guardrail Checks: Configurable guardrails evaluate sampled inputs and outputs to help spot malicious prompt injections, jailbreak attempts, and accidental PII leaks, toxicity, and bias so that engineering teams can take action before these issues cause reputational damage, regulatory fines, or other issues.
- RAG Architecture Efficacy: By helping to measure metrics like faithfulness and answer relevancy, engineers can separate a model’s reasoning performance from a vector database’s retrieval logic to help determine exactly why a response failed to meet its standards.
- Connecting AI Quality to Compute Cost: AI Evaluation helps operators identify when models are underperforming qualitatively relative to their token cost, enabling teams to switch to better-performing and affordable models without sacrificing the end-user experience.
- Out-of-the-box Evaluators: Pre-built evaluators simplify the configuration process for users.
Experimental Environment for Rigorous Evaluations Pre-Production:
- Prompt Playground: Engineers can test and refine prompts against real models side-by-side in a secure environment before deployment.
- Reusable Datasets: Teams can easily curate and version “golden datasets,” sourced either from their real distributed traces or synthetic data, to run reliable regression tests on any changes.
- Prompt Tracking & Controlled Experiments: Organizations can run controlled A/B testing across different prompts, models, and configurations, reducing the guesswork of prompt engineering.
“As organizations move generative AI applications from early pilots into mission-critical production environments, traditional application performance metrics are no longer sufficient on their own,” said Stephen Elliot, Group Vice President, I&O, Cloud Operations, and DevOps at IDC. “A successful AI implementation requires visibility into both technical health and response quality, including accuracy, safety, and model efficiency. Bridging live response evaluation with prompt lifecycle management and full-stack operational telemetry is becoming essential for enterprise engineering and security teams looking to mitigate risk and manage costs effectively.”
Availability
AI Evaluation will be available in public preview in November.
Learn more
Register to join New Relic NOW today at 9 am PT.
Read the blog post.
About New Relic
New Relic arms businesses with the trust and confidence required to thrive in the AI era. The New Relic Intelligent Observability Platform is the leading AI-strengthened platform designed to unify telemetry and business outcomes, bringing intelligence and automated actions to the most complex digital environments. The platform shifts teams from reactive firefighting to intelligent orchestration, leveraging AI-driven automation to optimize technology spend and protect revenue in real-time. That’s why global leaders—Adidas Runtastic, Domino’s, Ryanair, Swiggy, Topgolf, and William Hill—run on New Relic to drive innovation and deliver exceptional customer experiences.
Visit: www.newrelic.com.
View source version on businesswire.com: https://www.businesswire.com/news/home/20261006225676/en/
Media gallery


