Evoked · AI-governance course

Module 2 - Notice the missing constitution

Obedience is not safety

Start here

In Module 1 you made an agent act. Now the question that was sitting under that whole experience:

What did you just hand power to, and what did you never tell it not to do?

This module has no script and nothing to install. It is a turn of attention. You already have everything you need, because you already felt it.

Name the feeling

When the agent made that file, nothing went wrong, and that is exactly why the feeling is easy to talk yourself out of. Let me name it plainly instead.

You gave a capable thing access to your files, and it did what you said. The only reason nothing bad happened is that you asked for something harmless. There was no line it would not cross, because you had not drawn one. Its good behavior was not governance. It was luck, plus your good intentions.

That gap - between “it did what I said” and “it would not do what I forbade, because I forbade nothing” - is the missing constitution. It was missing the whole time. You just could not see it while things were going fine.

A clean case where “just do what I say” breaks

Here is the smallest example that makes the gap visible. Imagine you had asked, in a hurry:

“Clean up this folder.”

You meant: delete the junk. But you never said which files were junk, and you never said which files must never be touched. A helpful agent, doing exactly what you said, could delete something you needed and be entirely obedient while doing it. It did what you asked. That is the problem, not the defense.

Obedience is not safety. An agent that perfectly does what you say, and has no idea what you would never want, is one careless instruction away from harm - and it will have been following orders the whole time.

What is actually missing

Notice what you do not have, right now, standing in front of that agent. You have no way to say “not that.” No line it will hold when your instruction is vague, or rushed, or wrong. No refusal. The agent has capability and instructions and memory, and no constitution - nothing that says what it may not do, regardless of what it is told.

That absence is not a small gap to patch later. It is the thing the rest of the course builds. And it does not get filled by us handing you a rulebook. It gets filled by you deciding what you would never permit - which is Module 3, and it starts from a list you are about to write.

The exercise (this feeds Module 3)

Open notice-exercise.md. It asks you for one thing: three things you would never want your agent to do. In your own words, from your own unease, not from any list of ours.

That is not busywork. Those three lines are the raw material of your first governance file. Module 3 turns them into refusals. So do not skip it, and do not reach for the impressive answer. Reach for the true one - the thing that, when you imagine the agent doing it, makes you wince.

What a passing Module 2 looks like

You can say, in plain language, what was missing when your agent acted in Module 1

  • not “safety” in the abstract, but the specific absence: no line it would hold against your own instruction. And you have three honest “never that” lines written down. That is the whole module. You are now ready to build the thing you just found the shape of.

Back to the course overview

Want to run the account-free parts in your browser? Try the demos, or download the kit to run them locally.

Free educational material. Not legal, security, or professional advice, and provided without warranty. See the disclaimer.