democratised ai: the case for and against letting everyone build

written by Sonia Nash

Every AI platform vendor will say it at some point: everyone can build AI tools, AI agents, workflows and applications thanks to vibe coding. And it is largely true; we now see people with no technical knowledge building “stuff”, with varying degrees of usefulness and sophistication. The dopamine kick and the sense of pride most people get from doing something they’ve never done before is real and reinforced by an LLM that will always tell you what a brilliant idea you’ve just had. A lot of us use self-built little tools in the private sphere built on the cheapest annual license you can get from Anthropic, OpenAI, Lovable or the likes. Most people don’t even get to spend additional tokens, meaning that they think you can build a fully fledge product in a weekend on a £80 a year subscription.

Now, ask a room of executives whether employees should be allowed to experiment with AI and build their own agents and you get two confident answers. One half says yes obviously motivated by productivity gains and FOMO. The other half says no, usually on the basis that you’re one bot away from a data breach and a regulator’s letter.

The case for opening up

Back in the day (read pre-2022), if you were a finance professional with a process problem that required a system change, your chances to get a fix/solution/change request approved were next to none. First you had to carve out a significant amount of time building the case for change, on top of your day job. They you had to seek approval from a number of stakeholders, then the request was buried for months in the “who’s budget is this coming out of” conversation. And on the rare occasion the stars aligned and you had ticked all the boxes, you then had to go through the dreaded IT prioritsation list, and your little fix was never critical enough to make it to the top of the pile. As some of you will know, this is an over simplification of the absolute nightmare this process really was, underpinned by the fact that most non-IT folks could not even test the solution they were proposing. By the time you were told by IT the fix couldn’t be put into production for whatever reason, you had lost both the will to live and the motivation to suggest change number 2.

Letting the people who actually know the process experiment is a testing mechanism before it is a delivery one. Because they live in that process day in day out, they are able to map friction points, dependencies, workarounds and build a solution within one afternoon. Even if half of what gets built is discarded, the organisation can quickly determine the viability of a solution without spending enormous amounts of work and additional consultancy costs. Two weeks of someone building a prototype or a rough agent resolves arguments that would otherwise take week. And it is honest in a way that specifications never are: for example, you discover that the master data is unusable in week one rather than at UAT.

Another benefit of opening up lies in practical literacy. Organisations with no hands-on experience get sold badly. They can’t tell a thin wrapper from a platform. They can’t challenge a vendor’s benchmark claim because they never ran an evaluation. They over-scope pilots, under-scope data preparation and then sign a three year license for capability that can be built in a fortnight. See practical experience as a procurement asset; you will have a better understanding of what it takes (and doesn’t take) to improve a specific process.

Finally, you will find a variety of sources showing that as much as 80% of workers use unapproved AI tools at work. The debate isn’t about experimentation versus no experimentation, it is experimentation you can see versus the one you cannot. Restrict access to AI tool and the activity will move to personal accounts where you have no logging, no data residency, no retention control and no idea what has been pasted into what. Approved experimentation converts invisible risk into a visible one and it is what will ultimately keep your business safe. You will have very capable and innovative people who want to do the right thing, therefore given them the right tools to do it will serve your interests.

The case against

Read the previous sentence again. Not all employees are capable and innovative when we talk about building applications/workflows. Not all of them want to either. Additionally, simply buying a tool and telling people “you have access to it, have fun” rarely generates viable solutions. Yes vibe coding makes building a lot easier, but without solid training, most people will not know where to start, let along building something truly transformative. Ask your non-IT employees what MCP or APIs are and you’ll quickly see what I mean. Unfortunately, most organisations do not provide adapted training, resulting in a workforce using expensive tools for very little return, which aside from wasting money is also counter-productive in the long run; users will get disillusioned pretty quickly with the tech and abandon it while they may have genuinely good ideas that could make a difference. Even when enthusiasm exists, very few programmes ever book all this work to a budget line and return remains anecdotal.

The failure mode isn’t just the experiment itself. Someone builds an agent to draft supplier communications and it’s useful. Three colleagues start using it. Six month later it is part of the process, it has no owner, no tests, no documentation, no runbook and the person who built it has left the business. This is the Excel macro problem all over again: finance teams have spent thirty years discovering that the critical spreadsheet is always the one with a macro nobody can explain, and then pray for it to never break.

Add to this your company silos. People not communicating. You will have forty versions of the same assistant or agent within weeks. Beyond duplication, without a register how do you even know how many agents are now running across the organisation when you have given 2,000 people access to create them? How do you put governance around what you cannot see? And on the cost side, unmonitored token spend can go out of control pretty quickly and with no cost centre allocation. With no owners, no accountability, no KPI measuring, how do you justify the rampant token cost? This is where the debate around ROI vs ROE becomes even more relevant (see the AI Reality Check Series episode 3). Many employees won’t know which model to choose, have no concept of token consumption and how to keep it low (i.e. choose the right model for the right task) and end up building agents where a simple workflow would have done the job perfectly well.

But of course, cost is only one side of a very problematic coin. The other side is risk. “It works” doesn’t mean “it’s done correctly”. Vibe coded often means unreviewed. In practice it means credentials pasted into code, endpoints with no authentication, little to no testing for things like prompt injection, nobody measuring the error rate (because there’s no evaluation set or baseline to start with) and critically, no thought given to segregation of duty. An agent built on one person’s credentials will happily surface content to colleagues that access controls were specifically designed to segregate. A finance agent connected to a shared drive doesn’t know that the folder three levels down contains the restructuring model for example. And most companies discover this specific flaw the moment someone asks the internal chatbot a question they shouldn’t be able to answer.

Finally, we have the accountability gap. When a customer receives the wrong answer from an agent an employee built on a Friday afternoon, who’s responsible? An employee isn’t a software vendor, they have no professional indemnity, no design authority behind them and no change control. The liability lands on the business and the business may not be able to point to any decision it made about that system, because it never made one. The business may not even know that agent existed.

Is there a middle ground? The case for governed experimentation

I’d argue there is one and it is not about compromising between the pros and cons. There is very little I’d compromise on in the list of risks we’ve been through on both camps. Rather, I’d look into a different design.

a) Select the right people and train them. As we mentioned at the beginning, not everyone will want to experiment and not everyone should. Assemble a team of volunteers from each business function and make sure they are trained on the rules of the game below:

b) Allow experimentation in a proper sandbox only

“Duh!” I hear you say. But most sandboxes fail for one reason: the data in them is totally useless so the work done in them is not representative of reality. Hence why many pilots never make it to production. What I mean by “proper” sandbox is:

  • Representative data, not real data. That is the one investment that makes everything else possible: a masked or synthetically generated extract with realistic messiness i.e. duplicate supplier names, inconsistent date formats, the 3% of records with a cost centre missing. A common mistake is creating clean synthetic data which will show you nothing of how the agent/solution will behave in the real world.

  • A pre-approved tool stack. A short list of models and platforms, not because the alternatives are bad or costly but because “which tools count as approved” is the question that stalls most experiments for two weeks.

  • No route to the outside. You’ve seen the news… We want no writes to production systems, no customer contact, no external publication outside of the sandbox. Easy to enforce technically and easy for people to understand.

  • Time limits. Sandbox work expires after, say 90 days unless it’s been promoted or renewed. You may already have regular sandbox refreshes as part of your processes, but make sure they are clearly communicated so everyone is on the same page.

c) Tier by consequence, not by technology

Yes I live for governance, but there is an element of criticality that must come into play. You don’t go into the same levels of scrutiny for an agent that triages marketing assets into folders versus one that interacts with your customers or suppliers, or takes actions into critical processes such as payments, or has access to personal, financial or commercially sensitive data.

d) The register: the spine of the whole thing

Everything else depends on it. Tiering means nothing if you can’t see what’s tiered, the promotion path in the next point has nothing to promote from, and the regulatory question is unanswerable. Make registration the route to access, not just a reporting tool. registration is how you get the API key, the sandbox data, the tool license. You are not asking people to document what they’ve done, you’re giving them what they want in exchange for two minutes of form filling. The register must include key elements which you’d expect to see: what does the agent do, who is the owner, what’s the tier, data it can access, systems it can write to, who uses it, decision it influence, date created etc

e) A promotion path AND a demotion path

Open experimentation in a sandbox with no route from prototype to production is a waste of everyone’s time. There has to be a path to promotion. 4 stages that may work in practice:

Experiment → Pilot. Someone other than the builder wnats to use it. It has to have a named owner, a defined user group and one measurable claim or outcome i.e. it saves the AP team around 6h a week. It doesn’t need to be proven just yet, but at least stated so that its value to the business can be assessed and we have something to test against.

Pilot → Supported. This requires a full evaluation, a security review appropriate to the tier, a runbook, an engineering owner who is not the original builder, an agreed benefit line with finance. In exchange the organisation takes on support, monitoring and the liability. Before anything crosses into supported use you’ll need:

  • A test set: 20 to 50 real cases with known correct answers, drawn from actual historical work. Building this takes a day an is the single most useful day anyone will spend on the project

  • A measured baseline. What does the existing human process actually achieve on those same cases? How long does it currently take, how many touchpoints, what’s the error rate? Without the baseline, how do you know if the agent is performing better or has improved the outcome dramatically?

  • An error profile. Not just how often it is wrong but how it is wrong. Is it quite obvious or does failure is hard to spot i.e. the number looks right but isn’t? understanding the error rate and how easy it is to spot is paramount before putting any system into production.

  • A stated threshold. Agreed before the test, not after. “we will adopt at 95% on this set with no silent failures” is a decision. “it look pretty good” isn’t an agreed decision.

Supported → Retired. Every supported agent has a review date and a documented kill-switch. Somebody knows how to turn it off and what breaks when they do.

Demotion. A stage often overlooked. If a supported agent’s error rate drifts, or if the process it served changes, the agent goes back one tier or is turned off entirely. Without this, the promotion is a one-way accumulator and by year 2 you’ll have a 2,000 legacy agents.

It may sound painful but it is the safest way to bring a genuinely good solution to production.

Where governed experimentation may fail

This process, like any other only works with people owning it and monitoring it. 4 areas worth keeping an eye on:

It becomes an approval board by stealth: if your green-tier use-cases, that is those that are not critical, start needing sign-off, or the register forms grows past a single page, you are at risk of killing enthusiasm and the good work within two quarters.

The register rots: entries with owners who have left, review dates in the past unactioned, statuses nobody has updated in a year. We don’t want the register to become an achive of things that used to be true. It needs to be actively monitored and maintained.

The promotion path has no capacity behind it: You need to be honest about your capacity and resources to bring the pilots into supported. If you haven’t got the resources to make it happen, this whole endeavour will have been wasted. Having an approved cadence upfront (i.e. we will bring a maximum of 4 pilots into supported per quarter) helps set expectations. Ultimately, if the above process is respected, there will be a clear ROE and ROI for each pilot which may in turn justify a resource investment in the engineering team.

The argument about whether to let employees experiment with AI has already bolted. In most cases, the building has started, the tools are already in use on people’s phones or personal accounts; the real question is whether any of it is visible to the organisation. That reframing is what makes the middle ground rather obvious in my opinion. Open experimentation without visibility produces a fleet nobody can account for. Restriction without visibility produces the exact same result, perhaps with worse data hygiene. Visibility is the common requirement here and it is cheaper to build than most people think.

So the test for your own business right now is simpler than the debate suggests. If someone asked you today how many AI agents are live in your organisation, could you answer and would you be right?

Next
Next

Your AI model strategy is a finance problem, not just a technology one