InsightFinder has secured $15 million in new funding. InsightFinder says the new funding will help them build and ship tools to find and fix failures in AI agents faster.
From infrastructure monitoring to AI agent observability
InsightFinder AI grew out of academic work that began more than a decade ago. The company moved into commercial products in 2016, applying machine learning to monitor and auto-diagnose IT infrastructure. Helen Gu, founder and chief executive and a computer science professor at North Carolina State University, said the business has now pivoted its focus to the newer problems created by AI agents inside enterprises.
Gu said AI agents have changed what teams need to monitor in their stacks. Systems used to be predictable in a certain way — servers, databases, queues. Now models, data flows and infrastructure interact in real time. That interaction creates failure modes that don't fit neatly into the old categories.
"In order to diagnose these AI model problems, you need to actually monitor and analyze the data, the model, and the infrastructure together," Helen Gu, founder and CEO of InsightFinder AI, said. "It's not always a model problem or a data problem; it's a combination.
Sometimes, it's simply your infrastructure."
The firm said it has been applying those lessons to a product set designed to detect drift, diagnose root causes and in some cases trigger remediation automatically.
Series B led by Yu Galaxy
The company announced a $15 million Series B round led by Yu Galaxy. InsightFinder has spent years polishing algorithms to sift signals from noise in complex systems, and the new capital is intended to scale its engineering and go-to-market efforts for what it calls AI agent reliability.
InsightFinder's work traces back to roughly 15 years of research, the company says. Gu previously worked for IBM and Google before founding the startup. InsightFinder's history in observability is part of its pitch: it argues that understanding and acting on the interplay between models, data and infrastructure demands a different approach than traditional monitoring alone.
How AI observability differs
Traditional observability aims to gather telemetry across layers and alert engineers when metrics cross thresholds. That's still useful. But InsightFinder and others say it's no longer enough once AI agents are chained into business processes and customer journeys.
AI agents can surface errors in subtle ways — a gradual model drift that only shows up when a particular cache is stale, or an infrastructure bottleneck that changes request latency and causes a model to time out. Those faults look like model failures on the surface. They can also look like data problems. And sometimes they're purely operational.
InsightFinder illustrated the point with a customer example: a major U.S. Credit card company noticed one of its fraud detection models had started to drift. Because InsightFinder was monitoring the broader infrastructure as well as the model, it traced the issue to outdated cache on some server nodes rather than flaws in the model itself.
Product and capability: Autonomous Reliability Insights
The startup has introduced a product called Autonomous Reliability Insights. According to the company, the tool offers a feedback loop that spans model development, evaluation and production. It aims to automate detection, diagnosis, remediation and prevention for AI agent workloads.
That's a big lift — the product now spans development, evaluation and live production. We call detection any time behavior or outputs drift past set thresholds, for example when a model's accuracy drops or outputs shift unexpectedly. Diagnosis means isolating whether the change is driven by data drift, model degradation, infrastructure faults or a mix. Remediation can be anything from rolling back a bad model to updating stale caches or rebalancing resources. Prevention focuses on surfacing risks before they hit customers.
Gu pushed back on a common misconception: many organisations assume AI observability is mainly about evaluating large language models during testing. Instead, she said, a complete observability platform has to provide continuous, end-to-end feedback across development and production so teams can act quickly when agents go off course.
Why enterprises care
Companies that have embraced AI agents — whether for customer service, fraud detection, supply-chain automation or internal tooling — face a new class of operational risk.
Those risks can affect revenue, compliance and customer trust. That makes tooling that can tie an output back to a chain of causes valuable.
Enterprises are already spending on observability and incident response. But InsightFinder's pitch is that observability for AI needs to be more tightly coupled to model evaluation and data monitoring, not just logs and metrics. If a model's false-positive rate doubles, for example, teams need to know whether the spike came from a shift in input data, a code change that altered preprocessing, or a degraded storage tier that truncated inputs.
Detecting that quickly can stop small problems turning into major outages. It can also cut the time engineers spend working through noisy alerts that don't explain root cause.
Market context and challenges
The observability market is crowded. Vendors range from specialist startups to major cloud providers that now offer native monitoring services. Many tools still focus on telemetry collection and dashboarding; others add anomaly detection and alerting powered by machine learning.
InsightFinder positions itself where model monitoring meets traditional operational observability. That means building simple plugs for logs and metrics and deeper connectors into model-training pipelines and data stores so teams can trace root causes. Building those integrations at scale, and doing so in a way that respects enterprise security and compliance needs, is a non-trivial engineering task.
Buyers will also weigh integration complexity and security concerns, especially in regulated industries.judge solutions on how much noise they cut, how precise their root-cause analysis is, and how well automated actions perform in live environments. A false remediation — say, rolling back a model that was actually healthy — could be costly. So vendors must balance automation with guardrails and human oversight.
Outlook for InsightFinder
The fresh funding gives InsightFinder runway to expand its product and partnerships. The company will likely need to prove its approach at scale across different industries and deployment patterns. Financial services may move faster here since fraud models and compliance needs make traceability compelling. Retail, logistics and customer service are other obvious sectors.
InsightFinder's academic roots and long-standing focus on anomaly detection give it a technical foundation. But commercial traction will hinge on integrations, customer success and the ability to demonstrate clear reductions in incident time and operational cost.
Yu Galaxy led the latest round. Helen Gu will be the public face of the company's technical argument — she blends academic credibility with prior experience at large tech firms.
Related Articles
"A sound AI observability platform should provide end-to-end feedback loop support covering the development, evaluation, and production stages," Helen Gu, founder and CEO of InsightFinder AI, said.
This article was created with AI assistance.