The Machine That Learned Without Being Taught

The Machine That Learned Without Being Taught

In 2016, a machine made a move that experts thought was almost impossible. It wasn’t programmed to make that move. It had discovered it on its own. And suddenly, humanity had a very different question to answer: what happens when a machine starts finding ideas that humans never taught it?


The room went quiet.

On one side of the board sat Lee Sedol, one of the greatest Go players of his generation.

Across from him was something that couldn’t think the way he did.

It couldn’t feel pressure.

It couldn’t get tired.

It had never spent its childhood studying Go.

It had no memories of losing a game.

No teacher had ever told it:

“This move is beautiful.”

The opponent was AlphaGo, an artificial intelligence system created by Google DeepMind.

And on the second day of the five-game match in Seoul, AlphaGo played a move that made experienced Go players stop and stare.

It was Move 37.

At first, some commentators thought it was a mistake.

The move was so unusual that it seemed almost unnatural.

But it wasn’t a mistake.

It was a glimpse of something much stranger.

The machine had found a move that human players rarely considered.

DeepMind later estimated that the probability of a professional player selecting such a move was roughly one in ten thousand.

And suddenly, the match wasn’t just about winning a board game anymore.

It was about something much bigger.

Could a machine discover ideas that humans had never taught it?


The Game Humans Thought Machines Couldn’t Master

To understand why AlphaGo mattered, you have to understand the game it was built to play.

Go is an ancient strategy game that originated in China more than 2,500 years ago.

Its rules are surprisingly simple.

Players place black and white stones on a board.

The objective is to surround territory.

That’s it.

But beneath those simple rules lies an enormous universe of possibilities.

A standard Go board contains a 19 × 19 grid, giving players an astonishing number of possible positions and sequences.

Chess had already fallen to computers decades earlier.

In 1997, IBM’s Deep Blue famously defeated world chess champion Garry Kasparov.

Many people assumed Go would eventually follow.

But Go was different.

Much harder.

The number of possible positions is vastly larger, and many moves depend on intuition, pattern recognition and long-term strategy rather than simply calculating a few moves ahead.

For years, the strongest computer Go programs were still nowhere near the world’s best professionals.

Then AlphaGo arrived.

And everything changed.

A human Go champion facing an artificial intelligence opponent

The Machine That Studied Human Players

AlphaGo wasn’t born knowing how to play Go.

Its creators built a system that combined deep neural networks with a powerful search method.

One part of the system learned to predict promising moves.

Another estimated the likely outcome of positions.

And AlphaGo also played games against itself, using reinforcement learning to improve.

The result was something very different from the old idea of a computer following a gigantic list of instructions.

Instead of being told every good move…

it learned patterns.

It learned relationships.

It learned what strong positions looked like.

And gradually, it became better.

In October 2015, AlphaGo made history.

It defeated professional Go player Fan Hui 5–0 in a full match.

It was the first time a computer program had defeated a professional human player in an even game of Go on a full-size board.

But the real test was still waiting.


Then Came Lee Sedol

Lee Sedol wasn’t just another professional.

He had won 18 world Go titles and was widely regarded as one of the greatest players of his generation.

If AlphaGo could beat him, the achievement would mean something very different.

So in March 2016, the two opponents met in Seoul.

Human versus machine.

Centuries of accumulated human intuition versus a neural network trained on games and self-play.

The match attracted enormous global attention.

More than 200 million people watched the contest worldwide, according to DeepMind.

And then the games began.


The Moment Everything Felt Different

AlphaGo won the first game.

Then the second.

But the second game contained the moment that would become one of the most famous in AI history.

Move 37.

The move looked strange.

It wasn’t the kind of move many human professionals would naturally choose.

Some observers initially thought AlphaGo had made an error.

Then the analysis changed.

The move wasn’t a mistake.

It was brilliant.

A Go board illustrating the unexpected Move 37 played by AlphaGo

AlphaGo had found a strategic possibility that human players rarely considered.

DeepMind later described Move 37 as having roughly a 1-in-10,000 probability of being selected by a human professional.

Lee Sedol eventually lost that game.

But something else had happened.

For generations, Go players had built their understanding of the game from other human players.

One generation taught the next.

Patterns became traditions.

Certain moves became “correct.”

Other moves became “wrong.”

And now a machine had entered the game without being constrained by centuries of human habits.

It wasn’t simply remembering what humans had done.

It was exploring what might work.


The Human Strikes Back

But the story wasn’t over.

Lee Sedol was still Lee Sedol.

And in the fourth game, he did something extraordinary.

AlphaGo had won three games.

Lee needed a victory.

Then, on Move 78, he played a move that caught AlphaGo off guard.

DeepMind later described the move as another extraordinarily unlikely choice.

It became known as “God’s Touch.”

And this time, the human won.

Lee Sedol had found a move that the machine hadn’t expected.

For one evening, the score belonged to humanity.

The final match ended 4–1 in AlphaGo’s favor.

But the fourth game became almost as important as AlphaGo’s victories.

Because it showed something beautiful:

Human creativity had not disappeared.

The machine could surprise a master.

And the master could surprise the machine.


Then the Machine Did Something Even Stranger

You might think that defeating Lee Sedol was the end of the story.

It wasn’t.

It was only the beginning.

In 2017, DeepMind introduced a new version:

AlphaGo Zero.

And this time, the rules were different.

The original AlphaGo had learned partly from human Go games.

AlphaGo Zero didn’t need that.

It started with the basic rules of the game.

Then it played against itself.

Again.

And again.

And again.

Millions of games.

It began with random play.

It made mistakes.

It learned from those mistakes.

It updated its neural network.

Then it played again.

The machine became its own teacher.

And what happened next was almost absurd.

After only three days of self-play training, AlphaGo Zero defeated the previously published version of AlphaGo — the same system that had defeated Lee Sedol — by 100 games to 0.

After 40 days of training, it became even stronger.

The machine had effectively taken the board away from humanity.

And learned by itself.


It Wasn’t Memorizing Our Answers

This distinction matters.

AlphaGo Zero wasn’t simply searching a database for the move a human professional had played before.

It was learning through reinforcement learning.

The system played games against itself.

It evaluated positions.

It adjusted its neural network.

Then it played again.

The process repeated millions of times.

Eventually, strategies began to emerge.

Some resembled human ideas.

Others were unconventional.

And some appeared to reveal possibilities that human Go theory had overlooked.

DeepMind described AlphaGo Zero as discovering new knowledge and developing unconventional strategies through self-play.

AlphaGo Zero learning Go through self-play

That was the truly unsettling part.

The machine wasn’t just getting better at following human knowledge.

It was generating knowledge of its own.


What Happens When There Is No Teacher?

This is where the story stops being about Go.

Imagine a child learning a game.

You teach them the rules.

You show them examples.

You correct their mistakes.

You tell them what experts consider a good move.

Eventually, they become skilled.

That’s traditional learning.

Now imagine something different.

You give another learner only the rules.

You don’t show them a single professional game.

You don’t tell them what experts believe.

You don’t explain strategy.

You simply say:

Play.

And the learner plays against itself millions of times.

Eventually, it becomes better than the experts.

That’s what made AlphaGo Zero so important.

It demonstrated a powerful idea:

Under the right conditions, an AI system can learn complex strategies without first being given human examples.


From Go to Chess

The idea didn’t stop with Go.

In 2017, DeepMind introduced AlphaZero, a successor that applied a similar approach to multiple games.

Chess.

Shogi.

Go.

Again, the system started from the rules and learned through self-play.

According to DeepMind’s published results, AlphaZero surpassed the strongest chess program used in its evaluation after about four hours of training, surpassed the strongest shogi program after about two hours, and surpassed AlphaGo Zero in Go after about 30 hours.

The significance wasn’t that a computer had become good at board games.

Computers had been good at games for years.

The significance was the method.

A general learning approach was becoming capable of mastering different environments without being given a giant collection of human strategies.


The Machine Had Started Looking for Its Own Ideas

This changed the way researchers thought about artificial intelligence.

Traditional computer programs often depend heavily on human-designed rules.

Humans tell the machine what features matter.

Humans design strategies.

Humans define shortcuts.

Humans provide examples.

Systems like AlphaGo Zero and AlphaZero demonstrated another possibility.

Give the system a goal.

Give it the rules.

Let it experiment.

And allow learning to emerge from repeated attempts.

This approach is part of a broader family of methods known as reinforcement learning.

The machine doesn’t receive a textbook explaining how to win.

Instead, it learns from consequences.

A move leads to a stronger position.

Another leads to defeat.

Over countless trials, the system adjusts.

And eventually, patterns emerge that even its creators may not have anticipated.


And That’s Where the Story Gets Uncomfortable

When AlphaGo played Move 37, people asked:

Was that creativity?

It’s a difficult question.

The machine wasn’t conscious in the human sense.

It wasn’t sitting there thinking:

“I have an original idea.”

There is no evidence that AlphaGo experienced inspiration, emotion or intention like a human player.

What it did have was something else:

the ability to discover an effective strategy that wasn’t explicitly programmed into it.

Whether we call that creativity is partly a philosophical question.

But the practical result was undeniable.

Human experts began studying the machine’s games to learn from them.

The teacher had become the student.

And the student had become a source of new ideas.


The Legacy Was Bigger Than a Board Game

In 2016, AlphaGo’s victory was a spectacular demonstration.

But the deeper legacy was the research direction it helped accelerate.

The same broad ideas behind reinforcement learning, neural networks and search have influenced later AI research.

DeepMind has described AlphaGo’s legacy as helping inspire a new era of AI systems, with successors such as AlphaZero and MuZero extending related ideas to other problems.

The dream was no longer simply:

“Can a computer beat a human at a game?”

The more interesting question became:

“Can machines learn useful strategies in problems where humans don’t already know the best answer?”

That is a much bigger question.

And it reaches far beyond a wooden board.


The Strange Beauty of the Story

There is something almost poetic about the way AlphaGo changed Go.

Humans spent thousands of years developing the game.

Generations studied it.

Masters created theories.

Teachers passed those theories to students.

Then a machine arrived.

It studied human knowledge.

It played against itself.

And eventually it began showing humans possibilities they hadn’t considered.

Human and artificial intelligence exploring new Go strategies

Instead of replacing human knowledge, it added another source of discovery.

The board became a conversation between two different kinds of intelligence.

One shaped by history.

The other shaped by computation.


The Question We Are Left With

AlphaGo didn’t prove that machines think like humans.

It didn’t prove that AI is conscious.

And it certainly didn’t mean computers suddenly understand the world the way we do.

But it demonstrated something important.

A machine can sometimes reach an answer without following the path humans expected.

And that changes the relationship between people and intelligent machines.

For centuries, humans were the ones who created knowledge and taught machines how to use it.

AlphaGo showed another possibility.

Perhaps future AI systems won’t only learn what we already know.

Perhaps they will help us discover things we haven’t figured out yet.


The Oddnova Takeaway

The most important move AlphaGo made wasn’t Move 37.

It wasn’t the victory over Lee Sedol.

It wasn’t even the 100–0 victory of AlphaGo Zero over its predecessor.

The most important move was the moment humanity realized that a machine could search a space of possibilities and return with an idea we hadn’t expected.

Not because someone had written that idea into its code.

But because the machine had found it.

A board covered with black and white stones became a window into something much larger.

The future of intelligence may not belong entirely to humans…

and it may not belong entirely to machines.

Perhaps the most interesting future is the one where we discover things together.

And that may have been the real game AlphaGo was playing all along.


Sources & Further Reading

  • DeepMind — AlphaGo and the Lee Sedol match.
  • Nature — Mastering the game of Go with deep neural networks and tree search.
  • DeepMind — AlphaGo Zero: Starting from scratch.
  • Nature — Learning to play Go from scratch.
  • DeepMind — AlphaZero and its application to chess, shogi and Go.

Leave a Reply

Your email address will not be published. Required fields are marked *