Sitemap

5 Rules for 2026 Enterprise AI

6 min readOct 30, 2025

According to the recent popular MIT report, 95% of AI projects fail.

But how they fail, and even the definition of failure, are worth exploring. Failure (or success) of an AI project, or any technology-related project, in an enterprise typically comes down to just a few factors. Here we explore 5 of them. You can use this framework as a sanity check during the planning, scoping, implementation, and deployment phases of your next AI project.

1. Design for Extensibility

Rigid designs are brittle by definition. Conversely, resilience breeds sustainability. Modern AI systems are agentic, enabling an AI “expert” to perform their skill and leave the rest to other agents, tools, and humans. This is not new. Agentic systems linked by agent-to-agent (A2A) or model context protocol for tool-calling (MCP) follow similar patterns as a microservices architecture: linked APIs infrastructure. This concept is rooted in the basic principle of object-oriented programming. Not only is this design extensible, but it can also be developed in parallel as parts and maintained with less risk of systemic disruption.

Extensible systems allow iteration through market learnings and can leverage the latest AI tooling. They can thereby enable better/faster/cheaper results. Over time, as 3rd-party capabilities advance to solve the most common/general problems, these systems offload the task of custom code management.

2. Empower vs. Replace

Buried in section 3.3 of the MIT study referenced above is the concept of “shadow AI”. What the researchers found was that while only 5% of centralized custom AI initiatives achieved positive economic results, in the same organizations, fully 90% of employees were using AI tooling in their own jobs.

Press enter or click to view image in full size

Source: MIT The GenAI Divide STATE OF AI IN BUSINESS, July 2025

How could this be? Similar to waves of technology revolutions, such as smartphones and desktop software, AI is brought into the enterprise front door by end users vs. the backdoor by IT. AI implementation across the organization is done without IT approval, as end users see the advantages of incorporating these technologies into their day-to-day work.

When AI builders think about AI empowerment, we often envision redefining the day-to-day work of employees through automation. This approach requires retraining and sometimes replacing the workforce. In fact, employees are ahead of the AI development teams, using general-purpose tooling like ChatGPT, Cursor.AI, Claude, Gemini and Copilot and other tools (or internal/private versions of the same). These employees are using AI to get their jobs done with higher quality and volume via Human+AI vs. prior human-only approaches.

3. Abstract your job (…then abstract again)

Despite these grassroots, sometimes unauthorized uses of AI, we are still in early days of AI adoption. Typically we seek support via AI to automate some portion of our existing workflows and tasks. As the capabilities of large language models expand, we too must expand how we think of the value we provide in that process.

In their recent paper, “Canaries in the Coal Mine, Six Facts about the Recent Employment Effects of Artificial Intelligence”, Chander and Chen of Stanford’s HAI institute explore the effects of AI on employment.

Press enter or click to view image in full size

Source: Canaries in the Coal Mine, Six Facts about the Recent Employment Effects of Artificial Intelligence”, Chander and Chen, August 2025

Here, significant differences appear between early career software developers vs. senior developers. The implication is that (as ever) the tasks at greatest risk of replacement by automation are repetitive tasks. If early career employees are more likely to do repetitive tasks than more senior employees, then we all must continue to “climb the hill” to identify a level of abstraction away from rote tasks. This is where the creative human mind can add the greatest value.

As an example, recently, while working on TRUIFY.AI we were working to identify the AI policies that should be included in our automated compliance testing and remediation platform. After some brainstorming, we came up with 10 policies that should be included in our agentic platform. However, someone asked: why not let AI decide? Indeed, a general LLM model came up with 37 policies, across regions, including unpublished draft policies and the addressable markets for each to include…and then created the eval steps and templates to include in our platform to ensure coverage. In short: in whatever you’re doing ask: Could AI do this better? Can I generalize this task to something broader?

4. Build Guardrails (…but don’t use them)

Once we start to see promising results from LLMs to improve workflows, the question is asked: “how can we ensure the model doesn’t deviate and crash”? Someone with an enterprise engineering mindset will be familiar with the process of building test cases, testing harnesses/automation frameworks and continuous integration/continuous delivery (CI/CD) processes to enable ongoing iteration and improvement.

Evals can help us to avoid inaccurate or misleading results from AI systems through continuous improvement. These checkpoints are critical in developing smart, responsible AI. Evals validate when results are correct. By contrast, guardrails identify when results are incorrect and operate as hard boundaries that prevent movement beyond acceptable pathways.

Imagine driving a car on autopilot through a windy mountain road. While we are glad to see the guardrails that keep us from heading off a cliff, we are hopeful that we never touch these guardrails with our vehicle, which could (at best) result in damage and stop our travel. Worse still, the guardrail could fail to our peril.

In the same way, while the existence of AI guardrails is mandatory for production enterprise AI, frequent failures prevent AI systems from achieving the desired gain, and can reduce trust and therefore future investment along the way. Avoiding errors is better than correcting them.

5. Measure Everything

No enterprise, whether public, private, government or NGO operates outside some measurement of return-on-investment (ROI). Early AI experimentation by enterprise was often seen as research, with no immediate requirement of ROI, at least in economic terms. Enterprise literacy and readiness for future AI endeavors were virtuous enough goals.

However, going into 2026, earlier “learning” investments are coming under increased scrutiny around measurable impact on the business. They are increasingly tied to the key performance indicators (KPIs) and Objectives and Key Results (OKRs).

Initial ROI is critical for AI solutions to move beyond MVP to scaled production, however these metrics must not be used simply for validation, but rather as a vital feedback loop to define resource allocation and future areas of investment.

Often metrics are created that leave the reviewer asking “relative to what”? Enterprises, compared to new/young companies, typically have decades of operational/financial data available to serve as baselines. As such, comparable metric techniques such as A/B/N testing, Pre/Post comparisons and case/control studies can help us understand the effects of AI solutions within existing enterprise workflows where they are applied. Understanding the lift generated by the investment in the AI solutions and teams can ensure continuous investment and iteration.

Conclusion

Critical also is the recognition that as we build AI into these systems we are operating in a cycle. We are changing those systems and in-turn changing the training data upon which future model iteration is dependent. Such an approach can create a so-called “echo chamber”. This is where AI is unable to generate creative solutions. Instead, it is limited to historical contexts, which now include increasingly optimized, and therefore narrow, AI recommendations. Developing a counterfactual hypothesis– how things might have been with a different approach, is still the domain of human creativity. It’s our job to imagine these alternate realities, then test them iteratively, exploring new contexts where no known optimization yet exists.

About the author

Cameron Turner is founder and CEO of TRUIFY.AI, serving the US Fortune 500 with AI solutions. Drawing from 30 years experience in enterprise data, including the founding of two companies (ClickStream, acq. Microsoft in 2009 and Datorium, acq. Kin+Carta (LON:KCT) in 2021) Cameron brings a wealth of hands-on experience in leading teams to deploy solutions that uniquely fit the technology and culture of his clients. Cameron also serves as General Partner of Oxonian Ventures, a VC fund investing in Oxford University graduates, and volunteers with the Venture Studio at the Center for Entrepreneurial Studies at Stanford University. Cameron received his BA from Dartmouth College, MBA from Oxford University and MS in Statistics/AI from Stanford University.

ODSC - Open Data Science
ODSC - Open Data Science

Written by ODSC - Open Data Science

Our passion is bringing thousands of the best and brightest data scientists together under one roof for an incredible learning and networking experience.