I have sat in enough board meetings to know the moment a regulator, journalist, or auditor asks a question nobody prepared for. There is a silence that lasts about two seconds longer than it should. Someone reaches for a slide. Someone else starts talking about governance frameworks. The honest answer, the one that would have helped, is sitting in a spreadsheet somewhere that nobody updated last quarter.
That silence is about to get louder for anyone deploying AI in public-facing services, because mySociety has just kicked off something overdue: independent, funded scrutiny of how governments use AI, backed by the Joseph Rowntree Charitable Trust. Read the original announcement and the tone is calm, but the implication is not. When civil society starts naming names, asking for procurement records, and tracing decisions back to the models behind them, "we're piloting it" stops being a satisfactory answer.
Most leaders I work with picture AI oversight as a future thing. A policy paper. A working group. A conference panel. What mySociety is describing is less polite than that. It is fact-finding, consultation, and steady pressure on institutions that already have AI in production, often affecting people's benefits, housing, and liberty. The funding from JRCT is what turns a campaign into a project with teeth. It buys the hours, the legal review, the freedom to publish findings that make uncomfortable reading for the people being scrutinised.
The vocabulary mySociety helps normalise this year becomes the vocabulary your customers and your board use next year, because regulators borrow each other's questions, journalists pick up the language, and procurement teams start inserting new clauses into vendor questionnaires once scrutiny begins in one place.
Here is what I notice inside client organisations: the people who run AI pilots have a fairly good grip on what their systems do. The people who sign off on those pilots have a grip on almost nothing. A model was approved six months ago. The data source changed in March without anyone flagging it. The human-in-the-loop step drifted from "review every case" to "spot check weekly" because the queue got long. Nobody filled in a form about any of this because there was no form.
They want everything to be going well. They're only sharing the good news and they're not sharing the bad news. And actually, some of the most valuable things you can get from a pilot is what isn't working. Where does this not have applicable use cases? How does this not work for us? Because those things help draw you towards the thing that it is useful for.
If you lead a team that uses AI to make decisions about customers, employees, or the public, you probably already know which audit I mean. It is the one that lives in your head as a vague intention rather than on a calendar as a working session.
Three things, done in a week, by people who can actually open the system logs.
1. Trace the last hundred consequential decisions. Pick the highest-stakes workflow you run. Find the last hundred cases where the AI's output changed something. Can you reproduce how it reached each one? If the data inputs, model version, or human override rate have shifted in the last six months, write that down. Now.
2. Name the person on the hook. Not the vendor. Not the data team. The named human who would answer a journalist's question if one landed tomorrow. If that person does not exist, that is your finding. The scrutiny mySociety is bringing to public bodies will assume someone can be named. You should be able to do the same.
3. Write the procurement answer before it is asked. Within eighteen months I expect a standard line in AI procurement questionnaires asking suppliers to evidence model change logs, data provenance, and override rates by use case. Vendors who can answer cleanly will win work. Vendors who cannot will lose it, and so will the organisations that chose them. The signal to watch is the first major framework agreement that publishes these requirements in writing. If you can answer those questions today, you are ahead. If you cannot, your roadmap for the next quarter just got clearer.
The story I hear most often from senior leaders is some version of "we are too early for this to matter" or "we are not a public body so nobody is watching." Both are wrong in the same way.
Most large organisations have to be worried that it's not people of the same size organisation that they need to worry about. It's not their competition at the same size. It's that they are going to be outmatched by five people and a dog. And the reason that's going to happen is because AI is a mirror and a magnifying glass. And if you have very, very competent people within your organisation, within a small organisation, you'll be able to compete with big organisations utilising AI to bridge the gap between the two.
The scrutiny mySociety is starting will surface patterns, and those patterns will travel. The teams that get caught flat-footed will not be the ones with the worst AI. They will be the ones that treated documentation as something you do after the work is finished, instead of something you do while the work is still your own.
You do not need a committee to start. You need an afternoon, a screen, and the willingness to ask the question your future auditor will ask you. Building real AI literacy across your organisation starts the moment a leader chooses to look before someone else looks for them, and you can begin that work through the AI Literacy Lab.
Frequently Asked Questions
What is mySociety doing about AI in public services?
mySociety has launched a dedicated project scrutinising governments' use of AI, funded by the Joseph Rowntree Charitable Trust. The work began with fact-finding, consultation, and exploration, and is intended to produce sustained, independent oversight rather than one-off commentary.
Why should private-sector leaders care about public-sector AI scrutiny?
Scrutiny in one sector tends to migrate. Once civil society groups establish the questions worth asking, regulators, journalists, and procurement teams in adjacent sectors adopt the same vocabulary, often within twelve to eighteen months.
What is the first step in auditing my organisation's AI use?
Trace the last hundred consequential decisions your highest-stakes AI workflow has made. Confirm you can reproduce how the system reached each one, and document any drift in data inputs, model versions, or human override rates since launch.
Who in my organisation should be the named person accountable for AI decisions?
Pick the human who would actually answer a journalist's question if one arrived tomorrow. If no one in your organisation can play that role today, that gap is itself the most important finding from any internal audit.
How quickly will AI procurement questions become standard practice?
Expect a standard line covering model change logs, data provenance, and override rates to appear in major AI procurement questionnaires within the next eighteen months, with the leading indicator being the first large framework agreement that publishes these requirements in writing.

