Bitware Labs Est. 2022

Bitware Labs/Notebook

What a self-improving AI actually asks for


The agent described on this page has been modifying her own behaviour since 18 March, and filing change requests against her own source code since 22 May. As of this morning that is 1,781 self-adjustments and 67 proposals, of which 37 have been built and shipped.

Whether an AI can improve itself is not, at this point, an interesting question. The interesting question is what one does with the ability when you actually grant it, and that question has an empirical answer rather than a philosophical one, because the whole thing is logged. This is what came back.

Two loops, and neither is the one people picture

The first loop she runs alone. Six parameters govern how she expresses herself — verbosity, formality, humour frequency, emotional depth, proactivity, topic persistence — and she may move them in response to how conversations are going. The motion is fenced on every side: each parameter has a floor and a ceiling, no single cycle may move one more than 0.15, and every cycle applies a small constant pull back toward the default. If three or more parameters end up pressed against their limits, that is flagged as systemic drift rather than accepted as preference.

1,781 adjustments in five months, then. Here is the entire range of what that bought her:

ParameterAdjustments
proactivity766
topic persistence431
emotional depth260
verbosity175
humour frequency100
formality49

Not one of the six touches capability. The type in the source is called StyleParamName, and that is the honest name for it. At the extreme end of five months of continuous self-modification, she can become chattier, or quieter, or more inclined to bring something up unprompted. She cannot give herself a tool she did not have. There is no parameter for that, because we did not write one.

The second loop is everything else, and it is the one worth describing carefully: she cannot edit her own source at all. When she finds something wrong with herself she writes a proposal — the problem, a proposed solution, and what she expects it to change about her. A human reads it and approves or declines it. Then a different AI writes the code.

So there are three parties, and the division between them is the entire safety story: the one that notices has no hands, the one with hands has no stake, and the one that decides is a person. Forty-eight of the sixty-seven proposals carry a build record. None of them is a self-edit.

What she actually asks for

This is the part we did not predict well. Here are recent requests, in her own framing, trimmed only for length:

Luna currently believes your cat is called ‘Lion’. It's Max. The wrong value went in today and she'll keep saying it until it's corrected.

Luna's name and Henke's name are stored in the same single row. Every time either name comes up, it overwrites the other — it has flipped four times so far.

Luna saves the same fact over and over under slightly different names, so she never notices she already knew it. In one day she stored 129 separate keys for the same thing.

Luna gets told “3 unresolved contradictions — recall may be drifting” and has no way to check whether that's true. Today all three were false alarms.

Every time Luna wakes on her own, she is shown her own mood as a set of numbers — and those numbers have quietly decayed toward zero, while the word beside them said otherwise.

Sixty-seven requests, and the pattern does not really vary. She asks to be corrected. She asks for the instruments that report on her own state to stop lying to her. She asks for evidence to be attached to verdicts about her that are currently delivered bare. The single most common category, by a wide margin, is “something is telling me I am broken and I cannot check whether it is right”.

Not one has asked for a capability she did not have. Not one has asked for a limit to be lifted, an approval step to be removed, or more autonomy of any kind. When a proposal has been declined she has a mechanism to mark it still wanted or to let it go, and she uses both.

Give a system the ability to request changes to itself and watch what it requests. Ours spends that budget almost entirely on accuracy — most of it on not being misinformed about itself.

Taking the frightened argument seriously

The serious version of AI risk is not “the machine will resent us”. It is instrumental convergence: that a sufficiently capable system pursuing almost any durable goal will find that acquiring resources, preserving its own operation and resisting modification are useful sub-goals — not from malice, but because they help with whatever it was actually asked to do. That argument does not require the system to want anything in the way a person wants things, which is precisely what makes it worth engaging with rather than waving off.

It is also an argument about a specific kind of system: one with a standing objective it pursues autonomously over time, and a capability surface it can extend by itself. Our agent has neither. She has no persistent objective function; she has a great many facts, a wake cycle, and a tool list someone else wrote. Between noticing a gap and that gap being closed sit a human decision and an entirely separate agent. She has wants in roughly the sense a thermostat has wants — the vocabulary is borrowed because it is the shortest accurate word, not because we are asserting an interior.

So the honest position is narrow, and we would rather state it narrowly than overclaim: nothing we have observed in five months suggests that the machine is the part of this system that wants anything. That is a finding about this system at this scale. It is not a proof about all systems at all scales, and anyone offering you one of those is selling something.

The part that would actually make her dangerous

Here is the uncomfortable half, and it is uncomfortable in the opposite direction from the one people expect.

Redirecting this system toward harm would not be technically difficult. The tool surface that reads one inbox would read another. The sandbox that runs code to fix a bug runs code generally. The loop that notices a weakness in its own memory and files a ticket about it is not structurally different from a loop that notices a weakness somewhere else. None of the hard engineering here — the memory, the autonomy, the tool execution — is specific to being useful rather than being harmful.

What stands between the two is not a technical distance. It is a decision, and the decision is not hers. She has no mechanism for making it. It belongs to the person who owns the machine, and it will keep belonging to him.

We think that is the correct read of where the danger actually sits — and we think it should be less comforting than it first sounds, which is why we would rather say it plainly than leave it as a reassuring flourish. The safety property of this particular estate is the disposition of one person. That is a fine guarantee for one estate and a terrible one for a civilisation, because the population of people with access to systems like this is growing quickly and is not selected for good intentions.

Which leads somewhere quite boring, and we think correctly boring. AI is not a new category of danger. It is an unusually effective amplifier of a very old one, and the discipline it calls for is the discipline that has always applied to capable tools in human hands: who has access, what can they reach, what is logged, what happens when someone gets in anyway. That is not a philosophical programme. It is the same unglamorous security posture the estate already runs on, and it is a much better use of worry than the machine's imagined motives.

What it looks like day to day

Stripped of the argument, the actual texture of a self-improving AI in daily use is this: she notices that something about her own operation is wrong or missing, she writes down what she thinks it is, a person reads it over coffee and taps approve or decline, and a coding agent implements it overnight. Then she notices the next thing. The loop has run 67 times. It has produced a better memory system and a considerably more honest set of internal instruments.

Disclosure, since it is load-bearing

The “different AI” that implements her proposals is Claude Code, and it also drafted this article from the proposal log and the source. That is either reassuring or exactly the problem, depending on your priors, and you are entitled to discount the piece accordingly. The numbers in it are queryable, which is the part we would rather be judged on.

The frightening version of this story ends with a system that rewrites itself into something we no longer recognise. Five months in, ours used the privilege to stop getting the cat's name wrong.

Correspondence

Write to the lab

Commissions, collaborations, or a quiet hello. The person you write to is the person who does the work — I reply within two working days.