Artificial intelligence has quickly moved from experimentation to execution across financial institutions. What began as grassroots exploration of generative AI and small team pilots is evolving into a much bigger strategic question: What must a firm assemble, govern, and integrate to produce innovative and reliable business value with AI at scale?
That question is particularly urgent for business heads and tech decision makers given the current market is sending mixed signals. New model releases arrive constantly. Software vendors are repositioning existing products as AI platforms. Cloud providers are expanding their toolsets. Data platforms are embedding AI features. Financial data and workflow providers are exposing new interfaces for large language models, retrieval, and automation.
The result is a landscape that is concurrently full of possibility and difficult to navigate clearly.
In that context, firms that will extract durable business value from AI are those that have built the right architecture to connect intelligence to decisions—governed, traceable, and embedded in real workflows.
Enterprise AI is a stack of interoperating capabilities that must be selected, connected, and governed to drive real business outcomes, including models, APIs, orchestration layers, model context protocol (MCP) servers, agentic workflows, governed data, cloud infrastructure, monitoring, controls, and the business logic that ties all of those components to real work.
Each layer serves a distinct purpose, and no single layer is sufficient on its own. For example, a powerful language model without access to the right data will produce outputs that sound credible but are not trustworthy in production. Rich financial data without usable interfaces remains underutilized. Agents without governance frameworks create operational and reputational risk. Productivity tools without workflow integration generate short-lived curiosity rather than measurable, sustained business value.
The architecture itself is the differentiator, not any individual component within it. That’s a fundamental insight that separates firms making progress from those still running disconnected pilots.
At a practical level, enterprise AI architecture needs to address three distinct challenges:
Preparing data for AI use. Most organizations are managing fragmented internal data that is not consistently classified, tagged, or structured for AI workloads. Without that foundation, even the most capable model will produce unreliable outputs.
Managing agents in a controlled environment. As AI moves from answering questions to taking autonomous action across enterprise systems, firms need clear governance over which agents are registered, what data and tools they can access, when they are triggered, and how their activity is monitored, controlled, and auditable.
Redesigning workflows so AI can support meaningful outcomes. The most significant business value will come when firms rethink how work should happen once trusted data, agentic workflows, and human expertise operate together. Thinking beyond current processes is an operational differentiator.
Twelve months ago, the dominant AI use case in financial services was retrieval-augmented generation: Sophisticated search over structured and unstructured data, surfaced through a conversational interface. It is useful yet limited in its ability to drive fundamental workflow changes.
Today, task-level agents are live in production at a growing number of firms, including, for example, our FactSet AI for Banking and FactSet AI for Wealth. Multi-agent systems—specialized agents collaborating to complete complex, multi-step processes—are moving from demonstration to deployment. The next stage is always-on personalized agents that understand individual user preferences, analytical habits, and portfolio contexts to act proactively. It’s a near-term reality that multiple organizations are actively building toward.
The progression is significant because each stage raises the stakes for architectural readiness. RAG required good retrieval and data quality. Agentic workflows require all of that, plus governed tool access, event-driven triggering logic, audit trails, and controlled environments that can handle, and at times limit, autonomous action safely at scale.
Firms that invested in the underlying infrastructure early will find the transition to production agentic workflows much easier compared to firms that treated AI as a series of isolated product evaluations.
Trusted data is the foundation of enterprise AI, and the volume of financial data continues to compound at a rate that would have seemed implausible a decade ago. The dilemma is that more accessible data does not automatically produce better decisions. Without the right architecture beneath it, more data creates more complexity.
For AI workloads, the quality requirement is specific. Data must be cleaned, enriched, tagged with consistent metadata, resolved to stable entity identifiers, and structured in ways that allow AI systems to retrieve the right information under the right permissions at the right moment. Provenance must be preserved so every output can be traced to a verified source.
This is also where the distinction between generic AI demonstrations and genuine enterprise-grade implementations becomes clear. A language model running against web search and a language model running against governed, enriched, entity-resolved financial datasets will produce meaningfully different outputs on the same query. The former may look convincing. The latter is actually reliable.
To illustrate that concept, we handed Claude an equity analyst’s prompt: “I’m prepping for Apple’s next print. Give me the FY2026 consensus (EPS + revenue) and how it’s been revised, the split by product segment, where revenue actually comes from by country, who sits in the supply chain, and what’s driven the story this past quarter.”
We ran it twice: Once with only web search and once with the FactSet MCP connected. Everything below is the actual difference in what came back.
Source: FactSet
It's worth noting that the web-only output can vary with every query. The Claude + Web Search approach (left column) may surface different sources, different figures, and reach different conclusions each time. Because there is no governed data layer beneath it, the burden of verifying every source and every number falls entirely on the analyst, a process that is both time consuming and unsustainable at scale.
One of the most consequential developments in enterprise AI architecture over the past year is the emergence of model context protocol (MCP) as a standardization layer between AI systems and the tools, data sources, and workflows they need to access.
Before MCP, connecting AI to enterprise data and functionality required custom integrations for every model-tool pairing. That created friction, slowed adoption, and made it difficult to maintain a coherent architecture as models evolved.
MCP addresses that by providing a standardized way for AI systems to discover what tools and data are available, understand how to use them, and invoke them consistently and safely. The result is that firms and their vendors can expose capabilities through a common interface rather than rebuilding integrations each time.
The practical significance extends beyond technical convenience. For financial institutions, MCP enables AI systems to move from retrieval to action, from surfacing information to executing workflows across governed enterprise systems. An AI assistant can retrieve trusted data, invoke analytics engines, coordinate actions across internal platforms, and route outputs into review or reporting workflows, all while maintaining context, respecting permissions, and preserving traceability.
That capability class is a governed participant in the enterprise workflow rather than a conversational interface running over enterprise documents.
The move toward agentic AI requires up-front clarity on the operational requirements. Unlike a chatbot that a user initiates and controls, agents act autonomously, run continuously, access multiple systems, and trigger downstream processes. The concept of agent sprawl is already emerging in early enterprise deployments: Agents running unchecked, burning resources, creating audit-trail gaps, and producing outputs that conflict with each other or with firm policies.
The solution is to build a controlled and transparent agent-management environment from the outset, where agents are registered, their access is defined, their triggering logic is explicit, and their activity is visible and auditable across the enterprise.
Triggering frameworks are particularly important. The most efficient agents are those that know precisely when to act. Financial services firms already have deep experience with event-driven signaling such as market data changes, corporate actions, and earnings releases. That same logic applied to agent activation dramatically improves both efficiency and risk management.
The governance questions every firm should answer before scaling agentic workflows include:
Which agents are approved to operate, and what systems and data can they access?
What conditions trigger agent activity, and who can modify those conditions?
How are agent outputs reviewed, and which workflows require mandatory human oversight?
Can the firm reconstruct, for any given agent action, the request made, tools called, data accessed, output generated, and review completed?
What guardrails exist to ensure the agent’s actions remain aligned with its intended purpose?
These are the operational requirements that determine whether agentic AI can be trusted to support real financial workflows.
The buy-side research workflow illustrates concretely what the enterprise AI stack makes possible and what is at stake if that stack is poorly designed.
An analyst covering a public company following a quarterly earnings release may need to review the transcript, identify revenue drivers and margin commentary, compare management statements with prior quarters, validate key figures against structured financial data, pull historical estimates and peer comparisons, draft an internal summary, and prepare materials for the investment committee. The combination of structured data, unstructured content, historical context, and firm-specific analytical standards makes this a genuinely complex workflow.
A generic conversational AI tool can assist with drafting and summarization, but supporting a full research workflow without hallucinations requires access to governed structured financial datasets, entity-resolved content, permissioned internal research, entitled broker research and firm-specific applications with all of those elements connected through an architecture that preserves provenance so every output can be traced to its source.
When that architecture is in place, the analyst begins the process with a synthesized view of the most relevant developments. The system has already reviewed earnings materials, compared management commentary with prior quarters, validated key figures against trusted datasets, and produced a draft linked to supporting sources. The analyst's time shifts toward interpretation, judgment, and client engagement; all the work that requires human expertise.
That is the difference between AI that sounds useful in a demonstration and AI that functions reliably in production.
One of the most recurring strategic questions is where to build internal capabilities and where to rely on external platforms. The most useful framing is not a binary choice between building and buying, but a deliberate decision at each layer of the stack.
Few financial institutions will build foundation models; the investment is prohibitive and the capability is not where they create competitive advantage. The same logic applies to many other infrastructure layers where commercial offerings continue to mature quickly.
The differentiated opportunity sits in the layers closest to the firm's specific workflows, institutional knowledge, and client context. AI becomes more valuable when it can operate inside a firm's proprietary environment (rather than adjacent to it), working with the firm's internal research, client records, analytical standards, and operational processes.
A practical framework for most institutions is to:
Buy commodity infrastructure, cloud capabilities, and mature platform functionality.
Buy trusted external financial data and flexible workflow capabilities from established providers.
Build firm-specific workflows, business rules, and differentiated experiences where they generate meaningful competitive advantage.
Avoid committing to a narrow architecture too early, because the market is iterating quickly, costs change, model performance changes, governance tools improve, and standards mature.
To best position your firm for long-term success, build with enough openness to adapt and choose standards that preserve flexibility rather than constrain it.
AI pilots tend to generate enthusiasm because the experience feels novel and fast. For financial institutions, that is not a sufficient standard. A workflow that produces a polished summary or compelling output still needs to demonstrate that it improves execution, operates within appropriate controls, and scales reliably.
A practical measurement framework addresses three dimensions.
Business impact should be measured in workflow terms, not only technical terms. For research workflows, that means faster preparation for earnings events, reduced time gathering source materials, more consistent output quality, and greater analyst capacity for high-value work. For productivity use cases, it means time saved, faster knowledge retrieval, and reduced manual effort.
Risk and control effectiveness requires knowing whether outputs are accurate, supported, permissioned, and appropriate for the use case. Relevant measures include accuracy against verified sources, adherence to data entitlements, source coverage, data freshness, and the ability to preserve provenance back to underlying data and documents. For higher-risk workflows, firms should be able to reconstruct how any given output was produced.
Operational performance determines whether a workflow is sustainable at scale. Metrics include latency, uptime, tool-call success rates, cost per completed workflow, and the effort required to maintain prompts, integrations, and data pipelines. A system that produces strong results but is too expensive or fragile will not be ready for broader deployment.
The measurement cadence should be ongoing. Models improve or are replaced. Data coverage changes. User behavior evolves. Firms should establish a regular review process to assess whether AI-enabled workflows continue to perform as intended, and to retire low-value pilots while expanding those producing measurable returns.
In capital markets, intelligence that cannot be defended cannot be trusted. That standard applies directly to AI.
A model can be sophisticated, a response can be fast, and a recommendation can even be directionally correct. However, none of that matters if the result cannot be traced, validated, and reproduced within governed systems. Portfolio managers, research analysts, risk managers, compliance officers, auditors, and regulators all require the ability to understand not just what the AI produced, but why it is trustworthy.
For that reason, governance, lineage, transparency, and reproducibility are architectural requirements. AI systems operating in financial workflows must explain how they reached every conclusion and consistently reproduce the same result under the same conditions.
Financial institutions have spent decades building governed analytical environments for consistency, auditability, and trust. Enterprise AI must integrate within those environments rather than operate alongside them. When AI is embedded in governed financial systems rather than loosely connected to them, it becomes part of the institutional decision-making fabric rather than a tool that sits at its edge.
For decision makers assessing where to focus enterprise efforts, we recommend the following framework:
Audit your data estate thoroughly. Identify where your most valuable internal data lives, whether it is structured, tagged, and accessible for AI workloads, and where the gaps are. This is often more urgent than evaluating any specific AI tool.
Require auditability from the outset. Any AI solution deployed in a financial services context should be able to trace every output back to a verified source. If it cannot, the output is not production-ready regardless of how credible it appears.
Establish agent governance before you need it. The temptation is to run multiple agent pilots and worry about governance later. That approach enables sprawl and opacity problems that are harder to unwind than to prevent.
Choose infrastructure over individual tools. The specific models and applications available today will look different in a few short months. The data foundations, enrichment pipelines, and agent-management frameworks you build now will compound in value regardless of which models are leading at any given point.
Prioritize workflow redesign alongside technology deployment. The organizations capturing the most value from AI are pairing technology investment with genuine redesign of how work is organized and who is responsible for what.
Partner deliberately. The startup ecosystem in AI is producing remarkable capabilities at pace, but evaluating every new vendor would be a significant cost. The more sustainable approach is to work with a trusted partner that has already integrated the most relevant capabilities. Then build on a platform that allows the integrations to evolve without disruption to your core workflows.
The combination of exponential data growth, increasingly capable AI models, and agentic infrastructure is creating a permanent shift in how investment research, portfolio management, and client reporting will be delivered.
Organizations that approach this moment with intention and clarity about their data foundations, governance requirements, workflow priorities, and partnership strategy will find both today’s and tomorrow’s technology genuinely transformative rather than merely disruptive.
This blog post is for informational purposes only. The information contained in this blog post is not legal, tax, or investment advice. FactSet does not endorse or recommend any investments and assumes no liability for any consequence relating directly or indirectly to any action or inaction taken based on the information contained in this article.