◆ Independent guide ◆ Clear thinking ◆ Practical insights
Plain language

Nicole Junkerman on Reinforcement Learning, Plainly Put


A short, plain account of learning by trying and being rewarded, with no jargon, no algorithms and no claims about what happens next.

· 4 min read

Nicole Junkerman in a white top with her arms folded, standing in bright daylight beside a pale wall

Learning by trying

Most descriptions of how a machine learns involve showing it a very large number of correct answers. Reinforcement learning is the other kind. Nothing is shown the answer at all. Something takes an action, receives a signal about how that went, and adjusts. Then it does it again, and again, a great many times over, until the actions that tend to produce a better signal are the ones it reaches for.

The everyday comparison is learning a game nobody has explained. A child handed a controller with no instructions works out which buttons do something useful by pressing them and watching what happens. No rulebook is supplied. The score supplies the feedback, and the feedback supplies the rulebook eventually.

The reward is the whole design

The interesting part, and the genuinely difficult part, is deciding what counts as a good signal. Everything the system ends up doing is shaped by that one choice. Reward the wrong thing and you will get the wrong thing done extremely efficiently, which is a failure that has a long and faintly comic history behind it.

This is why the idea belongs in a general guide rather than only in a technical one. Asking what exactly is being rewarded here is a question anybody can put, and it is usually the right question well before anybody needs to know how the machinery works inside. It is the same question the governance checklist asks about purpose, in slightly different words.

Why it takes so many attempts

Trial and feedback is slow, because most attempts are poor. A few things follow from that, and they are worth knowing even at a comfortable distance from the work.

  • Early behaviour looks close to random, because it very nearly is
  • Progress arrives in uneven steps rather than along a smooth curve
  • Most of the practice happens in a simulation, where failing a great many times is cheap
  • Something trained under one set of conditions can do much worse under another

None of that is a criticism. It is simply what the method looks like from the outside, and knowing it makes the results easier to read sensibly.

Where the comparison stops

Learning by trying is a good way in and a poor place to stop. A child playing a game understands, at some level, that it is a game. Nothing in this kind of system understands anything of the sort. It is adjusting numbers in the direction of a better score, and the resemblance to the way a person picks up a skill is a resemblance rather than an explanation.

Which is also why this note makes no claims about where any of it goes next. The plain version is useful now: something tries, something scores it, and it tries again slightly differently. The plain-language glossary keeps the rest of the vocabulary in the same register, and the responsible AI pages take up the question of who is doing the scoring, which is usually where the interesting part lives. That is the whole of it, and Nicole Junkerman would rather leave it there than dress it up.

Keep reading across the guide

Move between the big picture and practical questions with Nicole Junkerman and the rest of this independent guide to AI in the UK.

Stay informed

New briefings, practical checklists, and city notes on AI in the UK, published here as the guide grows.

Read the latest briefings No sign-up needed, the guide is free to read.