Opsera Launches BrickForge, a Purpose-Built AI Operational Command Center for Enterprise Data Teams

Available on Databricks Marketplace

Evaluate your code base for modernization.

Reflections from three days at Gartner’s Application Innovation Summit, Las Vegas

This was my first time at Gartner’s Application Innovation Summit representing Opsera. The keynote framing set the tone immediately. Gartner’s opening thesis emphasized that the world is in a race to operationalize AI across applications, workflows, and development tools. But most technology leaders are struggling to scale deployments, drive real adoption, and show ROI without creating what they openly called “AI slop,” accelerating technical debt, or undermining user trust. The path forward, they argued, runs through human-AI partnership as a force multiplier, not through model capability alone.

Three days of sessions, roundtables, and conversations on the show floor confirmed a lot of what we’ve been building toward at Opsera and Forge. Here’s what I took away.

The sharpest idea didn’t come from a keynote

10x teams, not 10x developers

Before getting into what Gartner confirmed, the observation I kept returning to across the three days didn’t come from any single session. It surfaced in conversations: the goal should be 10x teams and systems, not 10x developers.

The enterprise AI narrative has been organized around individual acceleration: the developer who codes faster, the engineer who ships more. But individual speed without shared context, repeatable workflow, and governance discipline doesn’t compound across an organization. It breaks. The enterprises that will get real, sustained value from AI are the ones who’ve figured out how to extend context, consistency, and judgment across teams and make that the standard way of working. Basically, building well-oiled systems.

That framing is what makes everything else from this summit coherent. And it’s precisely what the human exponent in Gartner’s Business Value formula (which I’ll get to) means at the organizational level.

Gartner called out something we’ve been researching and building around

Your AI isn’t broken — your context is.

One session made a straightforward argument: your AI isn’t broken — your context is. In complex enterprise systems, when AI operates on incomplete architectural information, the problem isn’t just wrong outputs. Inconsistent outputs are harder to manage, because they erode the developer trust that adoption depends on. Once that trust breaks, it’s difficult to rebuild regardless of how the model improves.

Gartner addressed it directly saying AI can produce a confident answer and still be wrong. Push back, and it may acknowledge the error. That dynamic, confidence without reliability, is exactly what makes inconsistent outputs so damaging in enterprise environments. 

Judgment, analysis, and validation are the qualities AI still struggles to provide consistently, and they’re the qualities enterprise teams depend on most. The context layer is what closes that gap.

What made the context session land was what came alongside it. 

Gartner outlined six emerging development categories they expect to define enterprise software delivery. Context-driven engineering and spec-driven development were two of them. Looking at the full list, we’re already building across four of those six at Opsera and Forge, not because Gartner told us to, but because it’s what enterprise-grade AI development actually requires. That kind of external alignment matters, especially with enterprise buyers. Someone I spoke to on the show floor understood our token efficiency story immediately. The Gartner framing had already prepared them to see how context and spec-driven development reduce unnecessary generation.

The token rationing situation has a specific irony inside it

There was a roundtable specifically on AI token costs. The numbers being thrown around $100 to over $1,000 per developer per month under agentic coding workflows. The room was divided on how to respond.

The uncomfortable part is the sequence of events. Most of these organizations spent the last 18 months telling engineering teams to adopt AI, move faster, and use the tools. Now the same organizations are introducing token quotas, frustrating developers. Outputs are inconsistent because prompts are being constructed without the context layer that would make them reliable. And constant re-contexting is expensive. The token rationing is a symptom of the fact that governance and context architecture were not established before the adoption push, and now the cost of that gap is showing up on a finance report.

The formula for “Business Value’ from the opening keynote is worth writing down.

The keynote offered something concrete underneath the usual conference theme. Gartner’s framing for why AI programs succeed or fail:

Business Value = (Model capability × Workflow fit × Trust × Governance) ^ Human Exponential

This describes the failure pattern I see in enterprise AI programs more precisely than most explanations. The model is usually fine. What breaks is elsewhere: workflows that aren’t repeatable, trust that hasn’t been earned because outputs keep varying, governance that exists in a document but isn’t operational. When any one of those factors approaches zero, the value equation collapses regardless of how strong the model is.

The human exponent is the part that’s hardest to act on. Human judgment, adoption discipline, and change management are what determine whether the other factors multiply or cancel each other out. That’s not a soft observation, it’s the reason most AI programs fail to deliver at the scale organizations expect.

Gartner’s prediction on agents got a reaction in the room

70% of agents currently being built will fail!

One of the sharper claims from the summit: 70% of agents currently being built will fail. The reason cited wasn’t poor agent design, but the environment they’re built in that keeps shifting. Models change and behavior changes with them. Teams build agents like software prototypes, stand them up, ship the demo, move on, and no one maintains them as the underlying stack evolves. Token controls tighten, dormant agents keep consuming resources, and the surface area of what can go wrong expands constantly.

Building an agent is the straightforward part. Keeping it reliable as models update, policies change, and context drifts requires an operating discipline most teams haven’t built yet. It’s less an engineering problem than a governance and maintenance problem.

Gartner gave this gap a structural name: AIR — Action, Intelligence, Record. 

Most agent development over-invests in Action (executing tasks, generating outputs) while neglecting Intelligence (the situational awareness to operate reliably as conditions change) and Record (the governance, traceability, and context that makes an action auditable and repeatable). The agents in the surviving 30% treat all three as non-negotiable from the start. Record and Intelligence rarely get retrofitted successfully.

What good agent design actually requires

Alongside the failure prediction, Gartner was specific about what distinguishes agents that people trust and keep using from agents that get abandoned or quietly worked around. The qualities they identified:

  • Consistent, predictable agent behavior
  • Transparency about what an agent is doing and why
  • Feedback loops that let users understand and correct agent actions
  • Fairness and social appropriateness in how agents interact with people
  • User control over agent functionality, not just outputs

These aren’t UX refinements. They’re what the human exponent in the Business Value formula looks like at the design level. Organizations that build these qualities in from the beginning create agents that earn trust progressively. Those that skip them build agents that generate the kind of erratic confidence, correct sometimes, wrong confidently, that breaks developer trust in the first place.

Where the questions are now

The conversations at Gartner felt different from two years ago. The questions have shifted from “should we use AI” to “why isn’t it behaving consistently” and “how do we prove it’s doing the right thing at scale?” Those are harder questions. They’re also the ones worth spending serious time on, because the enterprises that work through them will build an operating capability around AI that compounds. That’s not something you can replicate just by switching tools. You need a structural layer above the tools that carries intent and context across sessions, teams, and the entire software lifecycle.

If you were at #GartnerAPPS last week, what were your takeaways? And if any of these questions live in your organization right now, let’s talk.

Evaluate your code base for modernization.

Recommended Content