Defining “the mess”

One of my inspirators, Dr. Russel L. Ackoff, wrote something along the lines of: “If you want to improve a system, you first have to define the mess”. This reminded me of something really important, the use of language and aiming to be precise as possible significantly matters to be able to communicate “a mess” clearly and unambiguously. This allows critical thinking to emerge and helps to build a consistent, mostly objective (even rational?) view of the situation at hand. However, some of that in the realm of AI has seemingly been lost.

The Antromorphization problem

I’m an avid user of AI in my daily work and personal life, so this is not bashing AI because I think it is in some ways magical and a great productivity boost in my field of work. I think of AI as systems, applications and machines. Things, not beings! The big problem I’ve been struggling with is the emerging use of terminology that disguises machine behaviour as human traits, applicable to beings. Words like: thinking, reasoning, intelligence, conciousness, awareness, state of mind, memories, etc.. I bet most of you have seen this all over your timeliness. To me, this blurs the line between humans and systems and in my opinion that is very unhealthy (to put it mildly).

No easy alternatives

On the other hand, I do understand the (commercial) appeal. Its inspiring to hear about development of intelligence and reasoning. Labeling these as fast, self-assessing word-guessing applications is a bit of a hard sell. This is of course combined with the fact that it is hard to define new fitting language for the system traits that are being developed under the umbrella of LLMs and AI systems. Our own language has evolved (and continues to evolve ) over millennia. This jump in progress, which it in my opinion is, has not had ample time to adapt our language and cultures properly.

Alignment

Alignment of systems (AI or deterministic) is essential. The software we build must always be aligned with human and organizational interest. The big advantage for deterministic systems is that they do not demonstrate emergent behaviour. AI systems change the game and that makes the alignment problem a continuous challenge.

Safety approaches

When we take a superficial look at the practices that are employed to counter the misalignment emerging in the actions of AI models, we see various guardrails employed. There is a tremendous amount of effort spend on grounding the systems, training data curation, pre-training, post-training, human reinforcement learning, etc.

A telling trait

We see these approaches have not been working satisfactory with the development of the latest frontier models. A piece of wisdom from the previous century informs us here on what is (not)happening. Peter Drucker said: “If you cannot measure it, you cannot manage it”. To me this translates into: “If you don’t understand it, you cannot manage it”. This applies to our current AI systems. With regards to understanding these, we seem to be in the “dark ages”, especially with regards to understanding how they work and how behaviour emerges.

Symptomatic

And so, what we see is that we’re now chasing symptoms rather than addressing the cause. Of course there are people already stating for a long while that the LLM architecture based on transformers is not the right architecture and therefore doomed to fail. There may be truth in that, but until now no evidence for a great alternative has been presented.

No Mistakes

A funny way that we’ve tried to use is of course, by adding better prompts, including the !Important “No Mistakes!”, but this is literally trying to add quality after the fact, it will never work.

Asking for alignment instead of taking control

Recently Anthropic announced their intent.md. This aims to create task specific alignment and feels like a band-aid that fixes a specific failure-mode.

In a similar way I was surprised to see that Microsoft announced a Code-of-conduct for Humanist AI. Although well-intended, I think this is asking an AI system to “play nice”. It does not exert or assume any control and responsibility over the potential outcomes.

The code of conduct problem

When I read the post/site of Microsoft in which they present their view on a code of conduct for Humanist AI, I instantly felt “something was off” by seeing the language and the framing of alignment creation. I do think the piece is a good starting point to contemplate on the idea of super-intelligence systems permeating the information World we know. We need to discuss what is happening and how we as a broad society see the rise of AI systems evolve responsibly.

On the other hand, this does not put the human front and center, but rather asks for AI systems to be considerate. Paragraph 2.2. makes clear that humans come last. The code-of-conduct is positioned as the most important control, but just as in the physical world, a code-of-conduct can be ignored and cannot be pro-actively enforced, which is a major problem.

The operator policies are interesting but also do not serve the purpose of the consumer and the broader society the AI system performs its work in.

Then as final point, a human can express preferences, which also can easily be dismissed or ignored. So, there is no value in expressing these besides having discussed what is desirable.

TL;DR; “Using a code-of-conduct to create alignment is a trainwreck waiting to happen”

Using a Code-of-conduct to discuss the control requirements of an AI system to protect our societies is a good starting point, if-and-only-if discussion is welcomed from all corners of the World. That also means this document needs to evolve before it can serve as global guidance.

Thought Police

“So then what?”, you might ask. Well… The terminology I have been using is “Zero-Tolerance policies”. When an AI system is “contemplating” to misbehave, it can be paused, then nudged to self-correct, if not, it needs to be shut down immediately. It is a bit like the “thinking police”, a digital ever-present monitoring system with an andon-cord to every AI agent/system instance. It can stop an agent mid-reasoning, interject a correction, and allows resume of operation if the system course-corrects. It sort of an omni-present force deciding on what can and cannot continue to run on a platform. The penalties that should result from misbehaving should add up, ultimately resulting in destroying the instance when mis-alignment is detected repeatedly.

We do this btw on deterministic systems, think of kubernetes with all observability tools, resource policies and system agents that stop, kill, restart pods when constraints are hit.

It is an effective way to manage machines like they are machines and applications, just that. No personality, nothing of enduring value lost, just boot a new one and continue.

On Assuming control

To wrap up the discussions, my take is that humans need to take control of the journey in which AI systems can support them to prosper. We are being sold a story where we need to stay passive and submit to the technology, assuming we cannot ever understand the systems we decide to create. That is detrimental to the human experience, and not in the best-interest to our society and cultures. That is not going to be easy, but better than to surrender to the creators of AI systems while we still are in the AI dark-ages.

Let me know what you think!