Splunk Is Where Cisco’s Enterprise AI Architecture Comes Into Focus
Three months after Cisco Live, Splunk’s new proof points show why inference location, model choice, data control, and token economics now belong in the same enterprise architecture conversation.
I enjoy going back to the promises made at Cisco Live once the announcements have had time to become products, integrations or something customers can actually test.
Three months after Las Vegas, it is probably not surprising that many of the clearest proof points are coming through Splunk. Cisco described the division of labor after Live as Cisco providing the critical infrastructure and Splunk serving as the intelligence layer for agentic operations. If that vision depends on agents having access to operational evidence, history and context, Splunk is where much of that information already lives.
At Live, I was still working through why Splunk had started to look like more than a security product. The announcements since then make its architectural role easier to see.
What interests me now is what happens to that vision as AI moves further into enterprise operations. A pilot can postpone questions that production cannot: Where does sensitive data remain? Where does inference run? Which model handles which job? What can an agent see and do? What does the operating model cost once token consumption and infrastructure become recurring expenses?
The announcements since Live do not answer all of those questions, but they make it easier to see where Cisco expects customers to encounter them. The location of inference is becoming an architecture decision rather than an implementation detail.
The Event SOCs Got There First
Some of the product promises from Live moved quickly. Machine Data Lake was in alpha, Agent Builder was headed toward a fall release, and AI Canvas was still part of the demonstrated roadmap. By August, Splunk said Machine Data Lake, Catalog and Agent Launchpad were generally available, along with expanded Federated Search and data-management capabilities.
Splunk also described Observability Cloud as one unified platform across applications, infrastructure, networks, digital experience, business processes and AI. That matters here because a decision about where inference runs eventually touches all of them. The model, the application using it, the infrastructure underneath it and the network connecting everything do not fail—or generate costs—in isolation.
The conference SOC work makes the progression easier to see. Cisco and Splunk used Amsterdam, RSAC, Cisco Live Americas and Black Hat to protect live networks while carrying integrations, detections and operating lessons from one event to the next. Cisco says the RSAC work informed the Agentic SOC at Live, and that detections developed at Black Hat would be used at Cisco GSX and in the first Agentic SOC at Splunk .conf, with the resulting searches, dashboards, playbooks and lessons carrying forward.
At Cisco Live Americas, the SOC included a UCS M8 with three NVIDIA GPUs to test local and protected AI workflows. Cisco's account says the team used AI Defense and DefenseClaw to inspect prompts, responses and agent actions around sensitive SOC data. It was still an event environment, not proof of a normal customer deployment, but it put a real version of the inference-location question inside the architecture.
The POD Is Not the Whole Story
Cisco AI POD for Splunk became generally available September 14. It is a pre-sized and validated on-premises deployment of Splunk AI Tier using Cisco UCS compute, NVIDIA acceleration, Cisco networking and Red Hat OpenShift. The initial release supports Splunk AI Assistant and Splunk AI Toolkit workloads; Agent Launchpad support on this self-managed layer is expected later this year.
The name is easy to misread. Cisco AI POD is not one universal appliance. AI PODs are modular infrastructure building blocks within Cisco Secure AI Factory with NVIDIA, and AI POD for Splunk is one workload-specific configuration.
Customers also do not have to buy the POD to run Splunk AI locally. Splunk offers AI Tier as a software path for qualified customer-provided GPU and Kubernetes infrastructure. That is the more important distinction for me. Cisco is creating a self-managed Splunk AI layer and offering the POD as its engineered infrastructure path, not making the hardware bundle a prerequisite.
The supported models add another choice. The current product page lists Cisco Deep Time Series Model, Gemma 4 31B Dense and OpenAI GPT-OSS 20B, with NVIDIA Nemotron open models expected later. These models do different jobs—the Cisco model is aimed at forecasting and anomaly detection—but the architecture is not tied to one external model provider.
None of that makes local inference automatically cheaper. The cost does not disappear when a model runs on infrastructure you control; it changes location. Model choice adds another variable to an operating model that already includes application performance, infrastructure utilization and data movement.
Splunk is beginning to instrument part of that problem. Tokenomics in Splunk Agent Observability tracks and attributes token spending across AI agents and employees' use of coding agents such as Claude Code, Codex and Cursor. Cisco says it will also forecast consumption patterns using the Deep Time Series Model.
I have not seen anything suggesting Splunk can look at a workload and say, “Run this one on your own GPUs instead of using an API.” But once an enterprise can measure consumption, observe model and application behavior, and choose where supported models run, I can see why it would eventually want those economics considered together.
Production Is Where the Architecture Shows Up
Local execution is not the only direction Cisco is pursuing. At .conf, Cisco said Splunk and AWS were expanding their relationship into joint product development through a multi-year agreement, with the work focused on security and advancing the Agentic SOC. Splunk Observability Cloud remains a cloud service, while Cisco Cloud Control was introduced at Live as the common environment for people and agents across the portfolio. This is not a disguised “everything comes back to the data center” argument.
Control also extends beyond where the model runs. Cisco completed its acquisition of WideField Security on July 31 and says the technology will be integrated into Splunk to correlate identity, session and activity telemetry across humans, non-human identities and AI agents. If an agent moves from recommending an infrastructure change to making one, the organization will need to know which agent acted, under whose authority and what happened afterward—the same decision-to-action boundary I explored earlier.
For years, hybrid architecture discussions were largely about where applications ran and where data lived. AI adds the location of the model doing the reasoning. The application can sit in one place, its operational evidence somewhere else, and the inference service somewhere else again.
Sometimes sending context to a cloud model will make perfect sense. Sometimes the data will be sensitive, regulated, expensive to move or simply something the organization has decided to keep under its own control. The answer will also change by workload; forecasting telemetry is not the same job as investigating a security incident or helping an operator write a search.
What I would ask Cisco now is how much of this architecture it is beginning to see in actual customer deployments. Not whether customers agree with the diagram from Live, but what they build after security requirements, sovereignty constraints, existing infrastructure, model fit and inference cost have each had a chance to change the plan.