Showing posts with label artificial intelligence. Show all posts
Showing posts with label artificial intelligence. Show all posts

Monday, January 19, 2015

Book Review: Superintelligence by Nick Bostrom

Superintelligence: Paths, Dangers, Strategies [Goodreads] by Nick Bostrom is a big idea book. The big idea is that the development of truly intelligent artificial intelligence is the most important issue that our generation will face. According to Bostrom, it may be the most important issue the human race has ever faced. This view suggests that how we approach the development and response to AI will be more important than how we respond to nuclear proliferation, climate change, continued warfare, sustainable agriculture, water management, and healthcare. That is a strong claim.
The sale of Bostrom's book has no doubt been helped by recent public comments by super entrepreneur Elon Musk and physicist Stephen Hawking. Musk, with conviction if not erudition, said
With artificial intelligence we are summoning the demon.  In all those stories where there’s the guy with the pentagram and the holy water, it’s like yeah he’s sure he can control the demon. Didn’t work out.
One almost wishes that Musk didn't live in California. He provided ten million US dollars to the Future of Life Institute to study the issue three months later. Bostrom is on the scientific advisory board of that body.
Hawking agrees with Musk and Bostrom, although without the B movie references, saying,
Success in creating AI would be the biggest event in human history. Unfortunately, it might also be the last, unless we learn how to avoid the risks.
Bostrom, Musk and Hawking make some interesting, and probably unfounded, presumptions. This is hardly uncommon in the current public conversation around strong AI. All seem to presume that we are building one or more intelligent machines, that these machines will probably evolve to be generally intelligent, that their own ideas for how to survive will radically differ from ours, and that they will be capable of self-evolution and self-reproduction
Jeff Hawkins provides the best answer to Elon Musk, Stephen Hawking, and Nick Bostrom that I have read to date:
Building intelligent machine is not the same as building self-replicating machines. There is no logical connection whatsoever between them. Neither brains nor computers can directory self-replicate, and brainlike memory systems will be no different. While one of the strengths of intelligent machines will be our ability to mass-produce them, that's a world apart from self-replication in the manner of bacteria and viruses. Self-replication does not require intelligence, and intelligence does not require self-replication. (On Intelligence [Goodreads], pp. 215)
Should we not clearly separate our concerns before we monger fear? The hidden presumptions of self-evolution and self-reproduction seem to be entirely within our control. Bostrom makes no mention of these presumptions, nor does he address their ramifications.
At least Bostrom is careful in his preface to admit his own ignorance, like any good academic. He seems honest in his self assessment:
Many of the points made in this book are probably wrong. It is also likely that there are considerations of critical importance that I fail to take into account, thereby invalidating some or all of my conclusions.
Beautifully, a footnote at the end of the first sentence reads, "I don't know which ones." It would be nice to see Fox News adopt such a strategy.
Another unstated presumption is that we are building individual machines based on models of our communal species. Humans may think of themselves as individuals, but we could not survive without each other, nor would there be much point in doing so.
We have not even begun to think about how this presumption will affect the machines we build. It is only in aggregate that we humans make our civilization. Some people are insane, or damaged, or dysfunctional, or badly deluded. Why should we not suspect that a machine built on the human model could not, indeed, would not, run the same risk? We should admit the possibility of our creating an intelligent machine that is delusional in the same way that we should admit the mass delusions of our religious brethren.
Is my supposition too harsh? Consider the case of Ohio bartender Michael Hoyt. Hoyt is not known to have had any birth defects, nor to have suffered physical injury. Yet he lost his job, and was eventually arrested by the FBI, after threatening the life of Speaker of the House John Boehner. Hoyt reportedly heard voices that told him Boehner was evil, or the Devil, or both. He suspected the Boehner was responsible for the Ebola outbreak in West Africa. He told police that he was Jesus Christ. Is Hoyt physically ill, or simply the victim of inappropriate associations in his cortex? We have many reasons to suspect the latter.
Bostrom originally spelled his name with an umlaut (Boström), as befits his Swedish heritage. He apparently dropped it at the same time as he started calling himself "Nick" in place of his birth name Niklas. Bostrom lives in the UK and is now a philosopher at St. Cross College, University of Oxford. Perhaps the Anglicization of his name is as much related to his physical location as the difficulty in convincing publishers and, until recently, the Internet Domain Name System, to consistently handle umlauts. His Web site at nickbostrom.com uses simple ASCII characters.
According to Bostrom, we have one advantage over the coming superintelligence. It is a bit unclear what that advantage is. The book's back jacket insists that "we get to make the first move." Bostrom's preface tells us that "we get to build the stuff." I tend to trust Bostrom's own words here over the publicist's, but think that both are valid perspectives. We have multiple advantages after all.
Another advantage is that we get to choose whether to combine the two orthogonal bits of functionality mentioned earlier, self-evolution and self-replication, with general intelligence. Just what the motivation would be for anyone to do so has yet to be explained by anyone. Bostrom makes weak noises about the defense community building robotic soldiers, or related weapons systems. He does not suggest that those goals would necessarily include self-evolution nor self-replication.
The publisher also informs us on the jacket that "the writing is so lucid that it somehow makes it all seem easy." Bostrom, again in his preface, disagrees. He says, "I have tried to make it an easy book to read, but I don't think I have quite succeeded." It is not a difficult read for a graduate in philosophy, but the general reader will occasionally wish a dictionary and Web browser close at hand. Bostrom's end notes do not include his supporting mathematics, but do helpfully point to academic journal articles that do. Of course, philosophic math is more useful to ensure that one understands an argument being made than in actually proving it.
Perhaps surprisingly, Bostrom makes scant mention of Isaac Asimov's famous Three Laws of Robotics, notionally designed to protect humanity from strong AI. This is probably because professional philosophers have known for some time that they are woefully insufficient. Bostrom notes that Asimov, a biochemistry professor during his long writing career, probably "formulated the laws in the first place precisely so that they would fail in interesting ways, providing fertile plot complications for his stories." (pp. 139)
To be utterly picayune, the book includes some evidence of poor editing, such as adjoining paragraphs that begin with the same sentence, and sloppy word order. I would have expected Oxford University Press to catch a bit more than they did.
Bostrom, perhaps at the insistence of his editors, pulled many philosophical asides into clearly delineated boxes that are printed with a darker background. Light readers can easily give them a miss. Those who are comfortable with the persnickety style of the professional philosopher will find them interesting.
Bostrom does manage to pet one of my particular peeves when he suggests in one such box that, "we could write a piece of code that would function as a detector that would look at the world model in our developing AI and designate the representational elements that correspond to the presence of a superintelligence... If we could create such a detector, we could then use it to define our AI's final values." The problem is that Bostrom doesn't understand the nature of complex code in general, nor the specific forms of AI code that might lead to a general intelligence.
There are already several forms of artificial intelligence where we simply do not understand how they work. We can train a neural network, but we cannot typically deconstruct the resulting weighted algorithm to figure out how a complex recognition task is performed. So-called "deep learning", which generally just means neural networks of more than historical complexity due to the application of more computing power, just exacerbates the problem of understanding. Ask a Google engineer exactly how their program recognizes a face, or a road, or a cat, and they will have no idea. This is equally true in Numenta's Cortical Learning Algorithm (CLA), and will be true of any eventual model of the human brain. Frankly, it is even true of any large software program that has entered maintenance failure, which is almost always an admission by a development team that the program has become too complex for them to reliably change. Bostrom's conception of software is at least as old as the Apple Newton. That is not a complement.
We will surely have less control over any form of future artificial intelligence than it will require to implement his proposed solution. Any solution will not be as simple as inserting a bit of code into a traditional procedural program.
Critically, Bostrom confuses the output of an AI system with its intelligence (pp. 200). This equivalence has been a persistent failure of philosophy. To quote Jeff Hawkins again, who I think sees this particularly clearly,
But intelligence is not just a matter of acting or behaving intelligently. Behavior is a manifestation of intelligence, but not the central characteristic or primary definition of being intelligent. A moment's reflection proves this: You can be intelligent just lying in the dark, thinking and understanding. Ignoring what goes on in your head and focusing instead on behavior has been a large impediment to understanding intelligence and building intelligent machines.
How will we know when a machine becomes intelligent? Alan Turing famously proposed the imitation game, now known as the Turing test, which suggested that we could only know by asking it and observing its behavior. Perhaps we can only know if it tells us without being programmed to do so. Philosophers like Bostrom will, no doubt, argue about this for a long time, in the same way they now argue whether humans are really intelligent. Whatever "really" means.
Bostrom's concluding chapter, "Crunch time", opens with a discussion of the top mathematics prize, the Fields Medal. Bostrom quotes a colleague who likes to say that a Fields Medal indicates that a recipient "was capable of accomplishing something important, and that he didn't." This trite (and insulting) conclusion is the basis for a classic philosophical ramble on whether our hypothetical mathematician actually invented something or whether he "merely" discovered something, and whether the discovery would eventually be made later by someone else. Bostrom makes an efficiency argument: A discovery speeds progress but does not define it. Why he saves this particular argument for his terminal chapter would be a mystery if he had something important to say about what we might do. Instead, he simply tells us to get on studying the problem.
I find that professional philosophers often slip in scale in this way. One moment they are discussing the capabilities and accomplishments of an individual human, generally assumed to be male, and the next they switch to a bird's eye view of our species as if the switch in perspective were justified mid-course. I find this both confusing and disingenuous. It is as if the philosopher cannot bear to view our species from the distance that might yield a more objective understanding.
The actions of individuals, both male and female, are inextricably linked to our cognitive biases. We do not make rational decisions, we make emotional ones, even when we try not to. We make decisions that keep our in-groups stable, by and large. A few, a very few, spend their days trying to think rationally, or exploring the ramifications of rational laws on our near future. A few dare to challenge conventional thinking aimed at in-group stability. Those few are not better than the rest. They are just an outward-looking minority evolved for the group's longer term survival. But the aggregate of our individual decisions looks much like a search algorithm. We explore physical spaces, new foods to eat, new places to be, new ways to raise families, new ways to defeat our enemies. Some work and some don't. Evolution is also a search algorithm, although a much slower one. Our species is where it is because our intelligence has explored more of our space faster and to greater effect. That is both our benefit and our challenge.
The strengths and weaknesses of the professional philosopher's toolbox are just not important to Bostrom's argument. Superintelligence would have been a stronger book if he has transcended them. Instead, it is a litany of just how far philosophy alone can take us, and a definition of where it fails.
I could find no discussion of the various types of approaches to AI, nor how they might play out differently. There are at least five, mostly mutually contradictory, types of AI. They are, in rough historical order:
  1. Logical symbol manipulation. This is the sort that has given us proof engines, and various forms of game players. It is also what traditionalists think of when they say "AI".
  2. Neural networks. Many problems in computer vision and other sort of pattern recognition problems have been solved this way.
  3. Auto-associative memories. This variation on neural networks uses feedback to allow recovery of a pattern when presented with only part of the pattern as input.
  4. Statistical, or "machine learning". These techniques use mathematical modeling to solve particular problems such as cleaning fuzzy images.
  5. Human brain emulation. Brain emulation may be used to predict future events based on past experiences.
Of these, and the handful of other less common approaches not mentioned, only human brain emulation is currently aiming to create a general artificial intelligence. Not only that, but few AI researchers actually think we are anywhere close to that goal. The popular media has represented a level of maturity that is not currently present.
The recent successes of the artificial intelligence community are a much longer way from general intelligence than one hears from news media, or even some starry eyed AI researchers. There are also good reasons not to worry even if we do manage to create intelligent machines.
Recent news-making successes in AI have been due to the scale of available computing. Examples include the ability for a program to learn to recognize cats in pictures, or to safely drive a car. These successes are impressive, but are wholly specific solutions to very particular problems. Not one of the researchers involved believes that those approaches will lead to a generally intelligent machine. These are tools and nothing but tools. Their output makes us better in the same way that the invention of the hammer or screwdriver, or general purpose computer, made us better. They will not, cannot, take over the world.
Bostrom is, at the end, pessimistic about our chances for survival. Perhaps this is what happens when one spends a lot of time studying global catastrophic risks. Bostrom and Milan M. Cirkovic previously edited a book of essays exploring just such risks in 2011 [Goodreads]. More information is available on the book's Web site. The first chapter is available online. These three paragraphs from Superintelligence anchor his position in relation to AI:
Before the prospect of an intelligence explosion, we humans are like small children playing with a bomb. Such is the mismatch between the power of our plaything and the immaturity of our conduct. Superintelligence is a challenge for which we are not ready now and will not be ready for a long time. We have little idea when the detonation will occur, though if we hold the device to our ear we can hear a faint ticking sound.
For a child with an undetonated bomb in its hands, a sensible thing to do would be to put it down gently, quickly back out of the room, and contact the nearest adult. Yet what we have here is not one child but many, each with access to an independent trigger mechanism. The chances that we will all find the sense to put down the dangerous stuff seem almost negligible. Some little idiot is bound to press the ignite button just to see what happens.
Nor can we attain safety by running away, for the blast of an intelligence explosion would bring down the entire firmament. Nor is there a grown up in sight.
One could imagine the same pessimistic argument being made about nuclear weapons. They must be reigned in before "some little idiot" gets his hands on one. Is that not what has happened? The Treaty on the Non-Proliferation of Nuclear Weapons has been a major force for slowing the spread of nuclear weapons in spite of the five countries that do not adhere to its principles. Separate agreements, threats, and sanctions has so far worked just well enough to plug the holes. Grown ups, from Albert Einstein to the current batch of Strategic Arms Limitation Talks negotiators, have come out of the woodwork when needed. Not only has no one dropped a nuclear bomb since the world came to know of their existence at Hiroshima and Nagasaki, but even the superpowers have willingly returned to small, and relatively low-tech ways of war.
Bostrom urges us to spend time and effort urgently to consider our response to the coming threat. He warns that we may not have the time we think we have. Nowhere does he presume that we will not choose our own destruction. "The universe is change;" said the Roman emperor Marcus Aurelius Antoninus, "our life is what our thoughts make it." Bostrom might learn to temper his pessimism with an understanding of how humans relate to existential threats. Only then do they seem to do the right thing. He might also observe that unexpected events should not be handled using old tools, as noted by the industrialist J. Paul Getty ("In times of rapid change, experience could be your worst enemy.") or management theorist Peter Drucker ("The greatest danger in times of turbulence is not the turbulence; it is to act with yesterday’s logic.") We will need new conceptual tools to handle a new intelligence.
Our erstwhile fear mongers seem also certain that any new general intelligence would, as humans are wont to do, wish to destroy a competing intelligence, us. People fear this not because this is what an artificial intelligence will necessarily be, but because that is what our form of intelligence is. Humans have always feared other humans and for good reason. As historian Ronald Wright noted in A Short History of Progress [Goodreads],
"[P]rehistory, like history, teaches us that the nice folk didn't win, that we are at best the heirs of many ruthless victories and at worst the heirs of genocide."
This raises the fascinating question of how we, as a species, would react to the presence of a newly competitive intelligence on our planet. History shows that we probably killed off the Neanderthals, as earlier human species killed off Homo Erectus and our earlier predecessors. We don't play well with others. Perhaps our own latent fears will insist on the killing off of a new, generally intelligent AI. We should consider this nasty habit of ours before we worry too much about how a hypothetical AI might feel about us. If an AI considers us a threat, should we really blame it? We probably will be a threat to its existence.
It is possible that a single hyper-intelligent machine might not even matter much in the wider course of human affairs. Just like the natural, generational genius does not always matter. The history of the human race seems to be more dominated by the slow, inexorable march of individual decisions than it is by the, often temporary, upheavals of the generational genius. How would human development have changed if the Persian commander Spithridates had succeeded in killing Alexander the Great at the Battle of the Granicus? He almost did. Spithridates' axe bounced off Alexander's armor. Much has been made of the details, but people would still spread through competition, and contact between East and West would still have eventually occurred. The difference between having a genius and not having a genius can be smaller than we think in the long run.
Bostrom's main point is that we should take the development of general artificial intelligence seriously and plan for its eventual regulation. That's fine, for what it is worth. It is not worth very much, really. We are much more likely to react once a threat emerges. That's what humanity does. Bostrom is at best early at delivering a warning and at worst barking up the wrong tree.

Sunday, October 19, 2014

Book Review: On Intelligence by Jeff Hawkins

On Intelligence [AmazonGoodreads] purports to explain human intelligence and point the way to a new approach toward artificial intelligence. It partially succeeds on the former and knocks it out of the park on the latter.
This is only book that Jeff Hawkins has written. Silicon Valley insiders may remember Hawkins as the creator of the PalmPilot back in the 1990s and, when the owners restricted his vision, he left to create Handspring. Both companies made a lot of money, which is all that matters on the Sand Hill Road side of Silicon Valley. The tech side of the Valley cares more about the fact that Hawkins succeeded in the handheld computing market where the legendary Steve Jobs had failed (with the Newton).
Hawkins' journalist co-author Sandra Blakeslee, on the other hand, has an Amazon author page that scrolls and scrolls.  She has co-authored ten books, several of which have related to the mind, consciousness and intelligence.  Her most recent book, Sleights of Mind: What the Neuroscience of Magic Reveals About Our Everyday Deceptions, was published as recently as 2011 with neuroscientists Stephen L. Macknik and Susana Martinez-Conde and was an international best seller. She has seemingly made a career out of helping scientists effectively communicate thought-provoking ideas.
Hawkins focuses all of his attention on uncovering the algorithm implemented by the human neocortex. Where that is impossible due to lack of agreement or basic science, he makes some (hopefully) reasonable assumptions and proceeds without slowing down. That will strike most neuroscientists as inexcusable. It makes perfect sense to an engineer.
Albert Einstein once said, "Scientists investigate that which already is; Engineers create that which has never been." Or, to quote myself, scientists look at the world and ask, "How does this work?". Engineers look at the world and say, "This sucks! How can we make it better?" There is a fundamental difference in philosophy required of scientists and engineers. Hawkins proved himself to be an engineer through and through even when he bends over backward when attempting to do some science.
There is a particularly useful review on Goodreads that drives a crowbar though the core of the book as if it were the left frontal lobe of Phineas Gage. The reviewer who goes solely by the name of Chrissy rightly points out Hawkins' overfocus on the neocortex.
It became clear that Hawkins was so fixated on the neocortex that he was willing to push aside contradictory evidence from subcortical structures to make his theory fit. I've seen this before, from neuroscientists who fall in love with a given brain region and begin seeing it as the root of all behaviour, increasingly neglecting the quite patent reality of an immensely distributed system.
Chrissy is correct. Hawkins' work is nevertheless critically important. Although the cortex is without doubt only part of the brain and only part of the "seat" of consciousness, his work to define a working theory of the "cortical learning algorithm" has lead directly to a new branch of machine learning. It is one that has borne substantial fruit since the book's 2004 debut.
It shouldn't surprise anyone that Hawkins' reviewers confuse science and engineering. Professionals are often confused on the separation themselves. Any such categorization is arbitrary and people have the flexibility to change their perspective, and thus their intent, on demand. To make matters worse, computer science is neither about computers nor science. It is the Holy Roman Empire of the engineering professions. Computer science involves the creation and implementation of highly and increasingly abstract algorithms to solve highly and increasingly abstract problems of information manipulation. It is certainly different from computer engineering, which actually does involve building computers, and it is also generally different from its subfield software engineering. Of course reporters and even scientists get confused.
Writing On Intelligence has not made Hawkins into a neuroscientist. That does not seem to have been his goal. Hawkins goal was to build a more intelligent computer program - one that "thinks" more like a human thinks. His explorations of the human brain have had that goal constantly in mind.
Hawkins himself states his goal differently, but I stand by my interpretation. Why? Consider what he says (pp. 90):
What has been lacking is putting these disparate bits and pieces into a coherent theoretical framework. This, I argue, has not been done before, and it is the goal of this book.
That makes him sound like a scientist. But he went on to do exactly what I claim. He described a framework and then implemented it as a computer program. That's engineering.
It seems almost strange that it took fully five years from the book's publication for Hawkins' group at the Redwood Neuroscience Institute (now called the Redwood Center for Theoretical Neuroscience at UC Berkeley) to publish a more technical white paper detailing the so-called cortical learning algorithm (CLA) described in the book. The white paper provides sufficient detail to create a computer program that works the way that Hawkins understands the human neocortex to work. Again surprisingly, another four years passed before an implementation of that algorithm became available for download by anyone interested. The Internet generally works faster than that when a good idea comes along. The only reasonable explanation is that a fairly small team has been working on it.
You can, since early 2013, download an implementation of the CLA yourself and run it on your own computer to solve problems that you give it. Programmers normally love this sort of thing. It is interesting to note that the Google self-driving car uses exactly the traditional artificial intelligence techniques that Hawkins denigrates in his first chapter. Hawkins may have come too late for easy acceptance of his ideas. There are entrenched interests in AI research and Moore's Law ensures that they can still find success with their existing approaches. A specialist might note that the machine learning algorithms in the Google car have stretched traditional neural networking well beyond its initial boundaries and toward many of the aspects described by Hawkins, without ever quite buying into his approach.
The implementation is called the Numenta Platform for Intelligent Computing (NuPIC). It is dual licensed under a commercial license and the GNU GPL v3 Open Source license. That means that you can use it for free or they will help you if you want to pay. You can choose.
Hawkins lists and briefs brief critiques for the major branches of artificial intelligence, specifically expert systems, neural networks, auto-associative memories and Bayesian networks. He is right to criticize all of them for not having looked more carefully at the brain's physical structure before jumping to simple algorithmic approaches. The closest of the lot is perhaps neural networks, which is notionally based on composing collections of software-implemented "neurons". These artificial neurons are rather gross simplifications of biological neurons and the networks, with their three-tier structure, are poor substitutes for the complex relationships known to exist in the brain of even the most primitive animals. Still, the timing of Hawkins book was unfortunate in that its publication occurred at the beginning of our current golden age of neuroscience. AI is back and AI research is suddenly well funded again. So-called deep learning networks currently contain many more than the three traditional layers, up to eight or even more. IBM has recently moved neural networks to hardware with their announcement of their SyNAPSE chip that "has one million neurons and 256 million synapses" implemented in silicon. All approaches are currently blooming for AI and are being applied to everything from voice and facial recognition to automatically filling spreadsheet cells to autonomous robots. There is currently less reason for the AI community to investigate, or lobby for hardware implementing, a brand new general approach. None of that makes Hawkins wrong. The human brain is still the only conscious system we know of and neuroscience is still doing a bad job of looking at its structures from the top down.
The largest single criticism of On Intelligence from me is that the cortex Hawkins describes is a blank slate, also called a tabula rasa. We know that the human brain is not. The idea that a mind is empty until filled solely by experience dates back at least to Aristotle. The Persian philosopher Ibn-Sīnā, popularly called Avicenna in Europe - a name still taught in Western universities, coined the term tabula rasa a thousand years ago as he interpreted and translated Aristotle's de Anima. We have known for decades that we are born with a number of innate functions, such as facial perception, so the brain is not a blank slate. Other animals have their own innate behavior such as the fear that many bird species have for the shape of a hawk. Hawkins does address the changing nature of brain function during life but does not even peripherally describe how innate functions fit into his theory.
Hawkins is often criticized for failing to provide a collated list of his assumptions. They are indeed buried in the prose. Hawkins comes right after the book's last chapter by providing an appendix that lists eleven predictions. They are all testable given the right science. Scientists are explicitly asked to validate or repudiate those predictions. A decade later, I am not aware of a comprehensive attempt to do so.
I have attempted to find all of Hawkins presumptions and have listed them here in the hope that they will both help other reviewers and neuroscientists who might pick away at them. All page numbers are from the 2004 St. Martin's Griffen paperback edition. All indications of emphasis are in the original text unless otherwise marked. The assumptions generally flow from the highest level of abstraction to the lowest, as Hawkins mostly does.
1. "We can assume that the human neocortex has a similar hierarchy [to a monkey cortex]" pp. 45. This one not only seems reasonable but is an assumption held by many scientists. It is in line with the many independent threads of evidence from evolutionary theory. Hawkins was intentionally careful when he used the word "similar".
2. "We don't even have to assume the cortex knows the difference between sensation and behavior, to the cortex they are both just patterns." pp. 100. This is actually a negative assumption in that he is not making one. This kind of thinking, determining what assumptions are necessary to a system, is in keeping with Hawkins' coding background. It is an engineering necessity.
3. "Prediction is not just one of the things your brain does. It is the primary function of the neocortex, and the foundation of intelligence." pp. 89. This is Hawkins' central idea and the one that informs not only the book and the implementation of NuPIC but the philosophic approach to his understanding of the brain and its functions. Hawkins relates the traditional AI approach of artificial auto-associative memories and declares, "We call this chain of memories thought, and although its path is not deterministic, we are not fully in control of it either." pp. 75. He proposes that "the brain uses circuits similar to an auto-associative memory to [recall memories]" pp. 31.
Here is also where Hawkins is forced to leave the cortex and venture into its relationships with another area of the brain. He notes the large number of connections between the cortex and the thalamus and the delay inherent in passing signals that way. He declares that the cortex-thalamus circuit is "exactly like the delayed feedback that lets auto-associative memory models learn sequences." pp. 146. He is onto something here, but one must question his oversimplification. The thalamus is also known to be involved in the regulation of sleep and thus almost assuredly implements more than just a delayed communication loop with the cortex.
Eventually he is able to bring his prediction model into sharp focus: "If the cortex saw your arm moving without the corresponding motor command, you would be surprised. The simplest way to interpret this would be to assume your brain first moves the arm and then predicts what it will see. I believe this is wrong. Instead I believe the cortex predicts seeing the arm, and this prediction is what causes the motor commands to make the prediction come true. You think first, which causes you to act to make your thoughts come true." pp. 102. This focus on the predictive nature of the neocortex is key to Hawkins understanding. Either the neocortex implements an algorithm really quite similar to the CLA as described by Hawkins and is therefore a "memory-prediction framework" or he has got it wrong. The predictive abilities of NuPIC suggest that he is on the right track in spite of his many assumptions.
4. Hawkins makes two interesting and useful assumptions for the purposes of developing a top down theory: "For now, let’s assume that a typical cortical area is the size of a small coin" pp. 138 (he does acknowledge there is substantial variation), and "I believe that a column is the basic unit of prediction" pp. 141. Why does it matter to Hawkins how large a cortical area is, much less a typical one? It shouldn't matter to a typical neuroscientist. They take the anatomy the way they find it. Remember though that Hawkins' purpose is to build a more intelligent computer program. He betrays his intent in making assumptions that all cortical regions have fundamentally the same structure (in spite of minor variations that he readily admits are in the literature) and in setting a typical size for an area of cortex. These assumptions will help him to design a computer program that learns in a new way. He is on better footing with the purpose of a cortical column. Cortical columns are indeed very regular in their construction and distribution, a fact that Hawkins dug out of 1970s research and relies upon heavily. It is striking and probably key to any successful high-level theory.
From this point forward Hawkins' assumptions get progressively more technical as he moves toward something that he can implement using existing technology. This may be the most important criticism of On Intelligence even though I personally find it perfectly excusable. Those seeking new neuroscience will be disappointed. Those seeking new and more general ways to approach artificial intelligence will be rapt.
Any review attempting to list Hawkins' more technical assumptions will need to pause to introduce new vocabulary for the general reader. A cortex, animal or human, is the outer layer of the brain. It consists of valleys and folds in order to increase its surface area in the small space afforded it in the skull. Its basic structure is a "cortical column" of six layers. The human brain has "some 100,000 neurons to a single cortical column and perhaps as many as 2 million columns." The Blue Brain Project of the Brain and Mind Institute of the École Polytechnique in Lausanne, Switzerland is currently attempting to model a complete brain, or at least the cortex. They have already succeeded in modeling a rat's cortical column. This is much more than Hawkins attempted, but a top-level theory of cortical function has yet to emerge from the project.
The six layers of a cortical column have many connections to other layers, other columns, other regions of the cortex and other areas of the brain. It is a complex network. Each layer consists of differently shaped cells. Hawkins collected the many, many neurons in a cortical column into functions at each layer. That alone may be a very valuable contribution if it is shown that level of abstraction can be made without sacrificing higher level function.
It will be useful and fascinating to see what emerges from a study of the Blue Brain Project's cortical column models. In the meantime, Hawkins has provided us with a roadmap of questions to ask.
5. Noting the obvious disparity between streams of sensory inputs and highly abstract thought, Hawkins illustrates how a hierarchical set of relationships between cortical areas could produce abstractions ("invariant representations") at the higher levels. "The transformation—from fast changing to slow changing and from spatially specific to spatially invariant—is well documented for vision. And although there is a smaller body of evidence to prove it, many neuroscientists believe you’d find the same thing happening in all the sensory areas of your cortex, not just in vision." pp. 114. Hawkins goes on to take this as written, which is just what he needs to do in the absence of established science in order to build a system.
6. Continuing with the vision system, possibly the best studied areas of the brain to date, Hawkins discusses some of the key regions called by neuroscientists V1, V2 and so on. He says, "I have come to believe that V1, V2, and V4 should not be viewed as single cortical regions. Rather, each is a collection of many smaller subregions." pp. 122. Hawkins is making a rather classic reductionist argument here. The question is not how arbitrary regions are defined or what they are called. The problem in front of our engineer is how they are connected. He needs that information to make reasonable (not necessarily physiologically accurate) assumptions if he is to uncover the mechanisms of the brain's learning system.
7. A region of cortex, says Hawkins, "has classified its input as activity in a set of columns." pp. 148. It is hard to argue with this suggestion given the success of Hawkins' artificial CLA in making predictions without the traditional training necessary to other forms of AI. Further, the cortex gets around limits on variation handling found in early artificial auto-associative memories, "partly by stacking auto-associative memories in a hierarchy and partly by using a sophisticated columnar architecture." pp. 164.
8. There are several assumptions about the detailed workings of a cortical column. "Let's also assume that one class of cells, called layer 2 cells, learns to stay on during learning sequences", says Hawkins (pp. 152). He makes no judgement whether that "learning" is innate or actively learned during life. He doesn't even know that it is really there. Something like it must be in order to make his theory work. That is no criticism! It is instead a testable hypothesis and thus the very model of scientific advancement. It also allows him to build something.
"Next, let’s assume there is another class of cells, layer 3b cells, which don’t fire when our column successfully predicts its input but do fire when it doesn’t predict its activity. A layer 3b cell represents an unexpected pattern. It fires when a column becomes active unexpectedly. It will fire every time a column becomes active prior to any learning. But as a column learns to predict its activity, the layer 3b cell becomes quiet." pp. 152. This might seem unjustified. What would make Hawkins jump to a conclusion in the apparently complete absence of supportive science. The answer is that the engineer clearly sees the necessity of feedback when it is presented to him. There simply must be a mechanism that fills the role or no learning could occur. Hawkins merely suggests a reasonable place for it and encourages the neuroscience community to look for it.
As for the lowest level, layer 6: "cells in layer 6 are where precise prediction occurs." pp. 201.
9. Finally, Hawkins rightly notes some differences between biological neurons and the artificial neurons used in neural networking models. It makes one wonder what IBM implemented on their SyNAPSE chip. How biologically correct were they? Hawkins says, "neurons behave differently from the way they do in the classic model. In fact, in recent years there has been a growing group of scientists who have proposed that synapses on distant, thin dendrites can play an active and highly specific role in cell firing. In these models, these distant synapses behave differently from synapses on thicker dendrites near the cell body. For example, if there were two synapses very close to each other on a thin dendrite, they would act as a 'coincidence detector.' That is, if both synapses received an input spike within a small window of time, they could exert a large effect on the cell even though they are far from the cell body. They could cause the cell body to generate a spike." pp. 163. This is exactly the sort of thing that can have great biologic effect and cause great trouble for overly simplistic implementors. It would seem that Hawkins was careful to avoid this over simplification even while embracing others.
Hawkins has also uncovered something really quite important and almost painfully subtle. Philosophers of mind, psychologists and priests have for centuries argued that the mind is fundamentally different from the body. We moderns have become comfortable with considering huge swaths of the body as mechanistic in nature. We can replace an arm, a leg, a kidney, even a heart for a while. We can insert a pacemaker, or a hearing aide. Surgery can cut, sew and sometimes almost magically repair, replace or augment much of our bodily infrastructure. We tend to view the body as a mechanism, however complicated, as a natural result. The brain, though, the mind, is a different matter. All the neuroscience conducted to date fails to convince most of us that the brain implements an algorithm. We cannot, so it is said, be reduced to an algorithm because that would imply that we could - one day - make a machine with all the abilities of people. Perhaps it would need to have all the rights, too. That scares people badly.
Parts of the brain have come to be accepted as algorithmic. Are you aware that a computerized cerebellum has been created for a rat? That was in 2011. Scientists and engineers are starting to soberly discuss creating such a device for paralyzed human beings.
The slow, painfully slow, admission that the body is a series of devices each of which chemically implement algorithms has been a long time coming. Parts of the brain have now unarguably fallen to the algorithmic worldview. First the ears, the eyes, the entire vision system. The cerebellum. The pineal gland. Hormonal balances. Most of the pons. Hawkins takes on the neocortex and, in spite of Chrissy's complaint, he did find it necessary to include the thalamus in his model. The bottom line is that the cortical learning algorithm is an algorithm. Philosophers of mind fear such a finding.
The idea that thinking is a form of computation dates from 1961 when Hilary Putnam first expressed it publicly. It has become known as the Computational Theory of Mind or CTM. Although CTM has its detractors (especially John Searle's Chinese Room, although that has been debunked to my personal satisfaction) it has become the basis for current thinking in evolutionary and cognitive psychology. The so called new synthesis of CTM is roughly a combination of the ideas of Charles Darwin's evolution, mathematician Alan Turing's universal computation and limits to computability proofs and linguist Noam Chomsky's rationalist epistemology. The basic idea is still the same, that human thought in human brains are algorithms even if they are quite complex ones that we haven't fully deconstructed. The new synthesis is about proving that theory.
"The dissociation between mind and matter in men and machines is very striking", observed David Berlinski in his book The Advent of the Algorithm[AmazonGoodreads], "it suggests that almost any stable and reliable organization of material objects can execute an algorithm and so come to command some form of intelligence."
We know what to do with algorithms. We implement them. It doesn't really matter how. We can implement algorithms in computer software or by creating DNA from a vat of chemicals or by lining up sticks and stones in clever ways. The only difference is the efficiency of the implemented algorithm. Electronic computers give us a way to perform calculations - implement algorithms - blindingly fast but they aren't the fastest way to implement all algorithms. Optical computers can do some things faster. Bodily chemistry, too. Or quantum computing. Each is just another way to implement algorithms be they designed by people or discovered by the search algorithm that we call evolution.
Discovering that the brain is algorithmic is arguably the most important realization of this or any other century. It means we can make more by any means we choose. That will shatter many world views even if Hawkins only got us part way there.