Live Trading News
Technology

The AI Safety Argument Is Being Held in the Wrong Room

The people who will actually be blamed are not in it. A jailbreak escapes a policy. The failure that costs an institution is an action that escapes an owner, and no amount of red-teaming a chat window touches it.

By Shayne Heffernan6 min readBullishVerified
Part of theKXCO Center
The AI Safety Argument Is Being Held in the Wrong Room

The people who will be blamed for what AI does are not in the room where its safety is argued.

The labs use the word jailbreak for a conversation that escapes a policy. Somebody finds a phrasing, a role play, a translation trick, and the model says what it was told not to say. That is a real failure and it will keep being real. Weights get stolen. System prompts get extracted. Safety training gets peeled off.

Institutions should be worried about a different failure, and they need a different word for it. An action that escapes an owner. A payment that goes out. A document that gets signed. A credential that gets issued. A position that gets booked. A record that gets altered quietly. Those are not sentences. They are facts in the world, with counterparties and statutes attached to them.

Confuse the two and you get what most boards have bought this year. Red-teaming for the chat window. An acceptable-use policy from counsel. The model wrapped in another layer of words. And at 18:40 on a Thursday the software still holds a credential that can do too much, and a tired person clicks yes.

Tokens are not authority

A language model produces tokens. Tokens are not authority.

Authority is a right, issued by a party with standing, to change something other parties will treat as real. Every bank already knows this in every other channel. A call-centre script is not a wire. An email is not a signature. A recommendation is not a booking. The confusion started when one interface began doing all four and the vendors called the result autonomy.

Autonomy in that sense means the software proceeds without a person in the room. That is precisely the property an institution should refuse to buy at full strength. What is worth buying is generation. Draft, retrieve, score, compare, surface the contradiction a reviewer would have taken hours to find. Then a governed step that decides whether the draft is allowed to become a fact.

The precise claim is not that jailbreaks are impossible. It is this. A jailbroken chat is still only information until it becomes a signed change made under a key that holds only the scope somebody issued. No bearer token. No implied administrator. No promotion from clerk to treasury because the model asked nicely. The model can beg. It cannot raise the ceiling and settle.

Exhibit 1. Two lanes, and nothing crosses between them.
Exhibit 1. Two lanes, and nothing crosses between them.

Exhibit 1. No tool accepts natural language as authority, so the jailbroken instruction has nowhere to land.

Intelligence got cheap. Judgment did not.

The fashion is to talk as though the remaining work is more intelligence. That is the wrong scarce resource.

A model can absorb, compress, retrieve and rehearse. It can look like memory and it can look like learning, which are two jobs a brain already does in tissue a surgeon can point at. Prefrontal circuits, basal ganglia, parietal maps and the neuromodulators all participate in choosing. Damage them and behaviour changes. None of that is in dispute.

What nobody has found is the part you could copy that is the responsibility. There is no gland of standing. There is no organ of being the person who can be blamed, licensed, struck off, imprisoned or forgiven. That remainder is not a feature request for the next training run. It is the reason a human still has to be in the loop.

Insight, compassion, intuition and experience do not ship in a model card. Anyone promising they will is selling standing they do not have. A machine can imitate the surface of those words well enough to fool somebody at the end of a long day. Imitation is not occupancy. A machine is not the party.

The objection worth taking seriously

The strongest argument against putting a human in the loop is that people comply.

Milgram put 26 of 40 subjects all the way to 450 volts. Jerry Burger's modern replication, published in American Psychologist in 2009 with proper ethics screening, still had 70 percent going past 150 volts against Milgram's 82.5 percent. Forty-six years moved it about 12 points. In Christopher Browning's account of Reserve Police Battalion 101, the commanding officer delivered the order in tears and openly offered any man an excusal. About 12 of 500 took it.

Anyone selling human oversight as a safeguard has to answer that, and most of them have never heard of it.

Exhibit 2. Who is being obeyed: a man in the room, against a written condition.
Exhibit 2. Who is being obeyed: a man in the room, against a written condition.

Exhibit 2. The Stanford Prison Experiment is left out on purpose. Le Texier showed in 2019 that the guards were coached.

Here is the answer. Milgram's subject is obeying a man in a room. Deference is doing the work, along with pressure and the ordinary wish not to make a scene in front of somebody in a lab coat. A written condition has no deference in it. It cannot be charmed, it does not get tired at 18:40, and it does not care who is asking or how senior they are. The person is answering to a process, and a process can be written down, versioned, and checked afterwards by a stranger who was never in the room.

That does not make anyone good. Run by a principal who intends harm, this produces well documented harm. Nobody should sell it as a conscience. What it does is keep the part that owes the world an answer inside a human being who can be named.

What to ask before you sign anything

Take these into the meeting. They work on every vendor selling an agent, mine included.

  1. Who owns the data the model reasons over, and can you leave with it.

  2. What is the exact scope of the agent's credential, in objects and verbs, not adjectives.

  3. Can the agent widen that scope by any route other than a person issuing it.

  4. Is there a preview of exactly what will change that does not pass through the model's own summary.

  5. Is refusing cheaper than approving.

  6. Are the approver and the executor separable for anything that leaves the building.

  7. Can a counterparty check a signature and a record without your server.

  8. What happens to the record if you are sold, shut or sanctioned.

  9. Which algorithms sign the record, and which certificate, if any, against which implementation.

  10. Where does content still leave into a rented model, and who accepted that residue.

A vendor who cannot answer those is selling a chat window with production credentials. A board that never asks has already appointed the tired human as its control system.

The line people quote at me

I have said we do not need more laws, we need better people. It gets quoted without the rest of it, so here is the rest of it.

It is not an argument against audit and it is not an argument against supervision. It is an argument against the belief that another statute will grow courage, restraint or taste in somebody who has none. Character is unfashionable because it cannot be procured. Rules without character become theatre. Character without a receipt becomes a story told after the damage.

Better people, then. Then infrastructure that assumes they will sometimes fail anyway, and makes the failure attributable to a name.

Words are not authority. The machine generates, the human decides, and the decision is signed so that nobody can move it later.

Shayne Heffernan, Ph.D., is the founder of Live Trading News, the KnightsBridge Group, Knightsbridge Law and the KXCO.ai ecosystem spanning post-quantum cryptography, identity, attestation and enterprise ontology.

Keep reading
AI agents

Your AI Agent and the Law: Who Acted, By What Right, and What You Can Prove

Your agent already acts in a legal system that was built for people. No new statute is needed for a payment it releases to be an act of your institution. Twelve months later a court asks who acted, by what right, and what can be reconstructed, and an export from a product you do not own is not a file. KXCO writes the act quantum-signed, time-stamped and marked on chain before the dispute. Knightsbridge Law sits on every Round Table ontology, so counsel is at the table before the act.

Shayne Heffernan10 min
$IBM

Quantum Is Accelerating

Quantum is accelerating. Not toward a machine that breaks RSA next quarter, which is still five orders of magnitude away on the first honest cross-platform yardstick the field has ever had, but toward foundries, clouds, logical qubits and government deadlines that are already fixed. Two clocks are running. Only one of them is slow, and it is not the one that decides what a bank, a court or a ministry has to do this year.

Shayne Heffernan23 min
$NVDA

Ignore the Noise: AI Is Not Slowing Down

Commentators pointed at a flat Nasdaq, a few missed prints, the cost of a megawatt and a safety letter, and called it a slowdown. It is not. Between 12 August and 16 September 2026 nine labs shipped frontier or near-frontier models, and the contest moved from one leaderboard to two stacks. What actually slowed is accountability. Enterprises still cannot put any of it into a court file. That gap, not the capability curve, stops deployment inside banks, hospitals, courts and ministries.

Shayne Heffernan22 min
post-quantum cryptography

Quantum Cybersecurity: The KXCO Chain Is Already Running

The World Economic Forum warned on 11 September that the quantum-safe race has changed gears. It is right about the direction and late about the work. Every KXCO product already runs on NIST's post-quantum algorithms, and all of them run on the same implementation: 42 package manifests declare one library. The chain verifies ML-DSA-65 on-chain at precompile 0x0b, tested live with three negative controls. NIST's own grader marked the library at 2,130 cases and zero failures.

Shayne Heffernan19 min
Read Live Trading News on Telegram

Every story, signed and delivered.

Subscribe to the kxco channel and get the headline, the AI-written key takeaways, and the chain-anchor link the moment we publish. Audio versions and per-ticker subscriptions arrive in the next iteration.

Open @KnightsbridgeInsightsNo email required.