AI

From Tool-Hopping to Talking: Reimagining Datacenter Observability with NETRA

Kaustubh Tripathi, Sreeman Repaka, Shejin Jacob Paulose28 September, 2026

URL copied to clipboard

It’s 3:00 AM. The glare of a monitor cuts through a pitch-black room as a critical alert pierces your sleep. Seconds later, you’re locked in a familiar, exhausting ritual: eight browser tabs open, bouncing frantically from Grafana metrics to Loki log streams, then over to Jira change logs. While the incident clock ticks, you’re forced to act as a human data aggregator – stitching together a story from a dozen noisy, fragmented telemetry sources.  

This is the “tool-hopping” trap: an engineering tax paid in high cognitive load, elevated MTTR, and operational burn-out.   

What if, instead of hunting through dashboards, you could simply ask your network what broke and get an exact, grounded answer in seconds?   

Enter NETRA (Network Reliability Admin),  an AI assistant built to eliminate operational friction by fusing real-time telemetry with curated documentation through a Model Context Protocol (MCP) fabric and a Retrieval-Augmented Generation (RAG) layer.

Before diving into how we engineered a solution, it is essential to look at why traditional telemetry systems break down under operational stress.

The Problem: The Cost of Fragmented Telemetry

Network operations at scale present a classical distributed data challenge. Telemetry is plentiful, but context is fragmented.

  • Lack of Insight: Most network tools only answer “what is the value of this metric right now,” providing raw data in graphs but leaving the cognitive work of interpretation and reasoning to the engineer.
  • Operational Fragmentation: During incidents, engineers are forced to manually switch between disconnected tools—such as Grafana for metrics, Loki for logs, Jira for change management and various network SOTs—to correlate data and “stitch a story together” under extreme time pressure.
  • On-call Bottlenecks: External engineering teams regularly reach out to on-call network personnel for routine status checks, creating avoidable operational friction.

Recognizing that raw metrics alone cannot solve cognitive overload during an outage, we fundamentally rethink our approach to infrastructure monitoring.

The Paradigm Shift: Dual-Channel Observability

NETRA isn’t just another dashboard, it’s a fundamental shift from monitoring to understanding. It shifts network operations from reactive dashboard monitoring to intent-driven understanding. Instead of querying individual database endpoints, engineers interact with a unified interface that balances two distinct operational layers.

Here is how we orchestrate that synthesis:

  • Live Signals (The ‘What’): For real-time state, NETRA queries volatile, high-cardinality data like telemetry, event logs, and operational dashboards. It’s the instant view of what’s happening in your network right now.
  • Curated Knowledge (The ‘Why’): When you need to understand topology, design, or procedure, the system taps into your organization’s stable knowledge base: runbooks, SOPs, and architecture references.
  • Deterministic Synthesis: Rather than returning raw database/monitoring links or unchecked summary text, NETRA cross-evaluates real-time signals against historical design intent to generate precise, cited responses.

Operational Modes

  • Proactive Monitoring via the Slack Bot: Operates directly in active channels leveraging MCP and RAG capabilities to handle automated alert fetching, correlation, multi-reviewer approval workflows, and Change Management Request (CMR) or activity publishing in real time.
  • Reactive Monitoring via the Web Application: Serves as an interactive interface where engineers can ask natural language questions, query live telemetry, and analyze system state using Model Context Protocol (MCP) clients, Retrieval-Augmented Generation (RAG), and integrated endpoints.

End-to-End NETRA Parent Architecture: Unifying proactive Slack bot workflows (event ingestion, LangGraph drafting, multi-reviewer approvals) and reactive Web Application interactions (MCP fabric, query routing, vector RAG) across live telemetry and curated knowledge bases.

At the core of both operational channels lies a shared foundation: a dual-retrieval spine designed to bridge volatile real-time states with stable, historical design intent. 

The Core Retrieval Engine: MCP and RAG Spines

Whether processing an automated alert in Slack or evaluating an interactive prompt in the Web Portal, NETRA relies on a unified core composed of two complementary technologies:

Standard AI implementations in infrastructure operations often fail because they rely on a single retrieval strategy. Pure RAG is blind to real-time state—you cannot vectorize millisecond time-series telemetry into a static database without constant, expensive re-indexing. 

Conversely, pure function-calling (tool use) lacks institutional context—fetching raw PromQL metrics or interface drop counts won’t tell an LLM whether a link is operating within expected design parameters or violating an internal SOP.   

To solve this, NETRA decouples retrieval into a Dual-Spine Engine: a live operational fabric (MCP) that queries volatile real-time state, and a semantic knowledge index (RAG) that retrieves historical design intent. Together, they ensure that every response is both real-time accurate and architecturally grounded

The MCP Advantage: Connecting the Live Estate

The Model Context Protocol (MCP) acts as a vendor-agnostic fabric that provides the assistant with a self-describing toolbox. It allows the model to interact with live network infrastructure such as observability tools, logs, and ticket systems in real-time, enabling agentic reasoning and multi-step workflows without needing custom integrations for every tool. 

Instead of writing custom API integration code for every telemetry system, NETRA configures independent MCP client instances per domain: 

  • Observability MCP Clients: Maintain persistent connections to time-series telemetry databases (Grafana/Prometheus) and log streams across Datacenter A and B.
  • Atlassian MCP Clients: Query active Jira operational tickets, scheduled maintenance windows, and change management records.
  • Netbox MCP Clients: Interface with infrastructure source-of-truth systems to query device inventories, IPAM data, and topology mappings.
  • Extensible Tooling: New datacenter pods or monitoring stacks can be introduced to the assistant simply by attaching a new MCP server endpoint to the routing mesh.

A full mesh: an independent, always-warm client for each datacenter’s live stack plus a client for change management and documentation. New capabilities arrive simply by attaching another server.

Grounding Truth: The RAG Knowledge Layer

While MCP handles the “now,” the Retrieval-Augmented Generation (RAG) layer manages the network’s design history. We chose RAG over fine-tuning because network documentation changes too rapidly for expensive model retraining; instead, we maintain a dynamic knowledge index. This layer processes everything from SOPs to architecture diagrams, using the BAAI/bge-small-en-v1.5 model for local embedding. By running this model internally, we ensure secure, reliable retrieval that captures semantic meaning without ever letting sensitive design documents leave our host or relying on external API availability. Using strict relevance floors, the system ensures that every answer is grounded in the organization’s written truth, serving live thumbnails through a guarded proxy to keep even visual data accessible and cited.

These vectors are stored in a dense-vector search cluster, allowing the assistant to find relevant passages at query time based on semantic similarity. To prevent hallucinations, matches must meet a strict relevance floor. The system even handles visual information; architecture diagrams are passed through vision models during ingestion to generate textual descriptions. When retrieved, the assistant serves a live thumbnail through a guarded proxy, ensuring that even complex topology is accessible and cited correctly. This entire layer runs as an isolated microservice, maintaining conversational uptime even during re-indexing.

Ingestion in four moves: split each document into overlapping passages, embed each with a local encoder, give it a content-derived fingerprint, and store it. Deterministic ids make the whole pipeline idempotent.

With this shared MCP and RAG engine providing full context, let’s look first at how it powers proactive operational workflows directly inside the channels where on-call engineers collaborate.

Netra on Slack: A Parallel Operations Interface

The end-to-end Slack interface pipeline: webhook ingestion, alert correlation, LLM-assisted drafting with RAG, and multi-reviewer DM approvals prior to channel publishing.

Alongside the primary web application, NETRA operates as a fully capable, parallel Slack interface integrated directly into active engineering channels. Serving directly inside operational and incident response workflows, the Slack integration brings MCP-driven real-time integration, automated reasoning, alert correlation, and RAG-powered context synthesis straight to where on-call engineers collaborate, enabling significantly faster and easier access to critical data.

Working hand-in-hand with the web platform, the Slack interface provides an end-to-end suite of capabilities spanning Activity/CMR publishing, critical-alert routing, real-time telemetry access, and interactive approval workflows:

  • Activity and CMR Publishing: Automatically formats and posts Change Management Request (CMR) notifications and maintenance updates across key lifecycle events. It ensures full operational visibility across teams while reducing manual communication overhead for engineers executing planned changes.
  • Critical-Alert Publishing Pipeline: Built to handle high-cardinality telemetry during major network events, this robust pipeline includes:
    • Event-Driven Ingestion: Listens to incoming webhook streams from Grafana, Loki, and internal monitoring systems, ingesting raw alerts with sub-second latency.
    • Deduplication & Correlation: Filters out alert storms by grouping related telemetry signals across devices, pods, and datacenters into a single consolidated incident payload.
    • Network-Aware Language Translation: Translates cryptic network alarm payloads and interface codes into human-readable operational summaries enriched with device topology context.
    • Incident RAG from Historical Threads: Leverages a specialized RAG index built by scraping historical Slack threads where engineers resolved past incidents, surfacing relevant troubleshooting steps and tribal knowledge directly inside the alert card.
    • Multi-Reviewer Approvals: Enables interactive multi-party sign-offs directly within Slack threads before executing high-risk mitigation actions or emergency changes.
    • Published Alert Behavior: Dynamically updates published alert cards in real-time as state changes occur, maintaining a single thread for the incident lifetime rather than posting redundant messages.
    • Reliability & Scale: Implements asynchronous queuing and rate-limiting controls to ensure uninterrupted message delivery even during severe multi-region outages.
  • Adjacent Activity Tracker Capabilities: Operates alongside incident handling to automatically capture operator actions, generate live timeline drafts, log troubleshooting commands, and draft post-incident review (PIR) documentation directly from active Slack thread interactions.

Working in tandem with the primary web application, this Slack interface ensures on-call engineers can seamlessly access live telemetry, incident context, and approval actions from whichever tool they are using.

Critical Network Alert notification  published to the external channel upon detecting a device down event in DC1.

Activity notification published to the external channel upon scheduling a network activity.

While proactive Slack workflows handle incoming alerts and team approvals in real time, complex ad-hoc investigations require an interactive diagnostic workbench.

Netra on the Web: A Reactive Web Interface

The life of one question: route, gather live evidence through the fabric, decide with the intent gate, add only-new curated context, then attach once, yielding a single grounded, cited answer.

The journey of a query begins with a safety inspection at the LLM Gateway, where the prompt undergoes rigorous authentication and policy validation. Once cleared, a primary router agent classifies the intent, determining the necessary datacenter scope and identifying the right tool domains. From there, the system enters an agentic loop within the MCP fabric. This isn’t a simple linear lookup; the model iteratively gathers live evidence, evaluates the results from one tool (like a Grafana metric), and uses that insight to decide which tool to pull from next (perhaps a specific Loki log stream). This loop continues until the model has sufficient context to resolve the ‘what’ of the situation.

Before finalizing the answer, an Intent Gate performs a critical check: does this query actually require design context? By filtering out requests that only care about transient live signals, we keep the experience crisp and noise-free. If the gate determines RAG is needed, the system retrieves relevant passages from the vector store, ensuring the documentation is complementary to the live evidence. This flow culminates in a centralized ‘attach-once’ formatting stage. Regardless of the complex internal paths the query took, this idempotent stage guarantees that citations and architectural supplements are woven in exactly once, delivering a unified, grounded answer that blends live metrics with architectural history.

Eg1. Netra in action: Checking health of secondary firewall prior to primary firewall upgrade

Eg2. Netra in action: Asking to check for any anomalies in the network which could have led to a blip in a given timeframe

When queried about complex infrastructure states, NETRA parses the request and formats the response into clean, scannable structures rather than wall-of-text paragraphs.

Giving an AI assistant deep visibility into live datacenter telemetry requires uncompromising operational guardrails. Intelligence is useless without trust—here is how we enforce strict boundaries across every interaction.

Security as a Foundation

Security is the foundation of trust in NETRA. All sessions are strictly authenticated, and sensitive credentials reside in isolated server-side environments. By restricting live network tools to read-only queries, we effectively bound the attack surface. This secure, read-only posture ensures that NETRA remains a reliable companion for network engineering, maintaining local isolation and data integrity across every interaction.

Every interaction is brokered by an internally hosted LLM Gateway, currently powered by our internal LLM. This gateway serves as a centralized governance point, enforcing strict security policies and guardrails. It filters sensitive parameters to prevent data leakage, ensures the assistant stays focused on its network engineering persona, and maintains a unified audit log for institutional safety compliance.

Engineering at the Speed of Conversation

By decoupling data collection from manual diagnostic reasoning, NETRA fundamentally alters the SRE operational workflow. By bridging the gap between volatile telemetry and stable design intent, we have effectively ended the era of reactive “tool-hopping.” Engineers are no longer forced to act as manual data aggregators under the 3:00 AM glare of fragmented dashboards. Instead, they can engage in a proactive, grounded conversation with the network itself. The infrastructure is no longer a silent black box to be watched and interpreted; it has become a partner that can be queried, providing automated context correlation and actionable understanding at the speed of thought.