Showing posts with label evolution. Show all posts
Showing posts with label evolution. Show all posts

Thursday, November 8, 2012

Steels - Experiments in Cultural Language Evolution

Luc Steels co-founded the Computer Science Department at Vrije Universiteit Brussel, and is part of their Artificial Intelligence Lab. In 1996 he founded the Sony Computer Science Laboratory in Paris, and is now  ICREA research professor at the Institute for Evolutionary Biology. Since about 1995, he has been heavily involved in finding practical ways to demonstrate how language evolves. His main approach, revealed in a wide range of publications, is to simulate the emergence of language in computational and, more recently, robotic agents. A recent book, Experiments in Cultural Language Evolution, published by John Benjamins, details these experiments by his team.  Here is an interview from 2006, courtesy of "Talking Robots." His work has resulted in a theory of language called "Fluid Construction Grammar" which reflects many issues brought up in usage-based approaches to language acquisition.


Review of  Experiments in Cultural Language Evolution

Reviewer:  Nick Moore
Book Title: Experiments in Cultural Language Evolution
Book Author: Luc Steels
Publisher: John Benjamins
Linguistic Field(s): Computational Linguistics; Historical Linguistics; Linguistic Theories; Text/Corpus Linguistics

SUMMARY

The ten papers collected in “Experiments in Cultural Language Evolution” represent the state-of-the-art of research into simulated multi-agent interaction. Centered around Luc Steels’ work at the Sony Computer Science Laboratory, Paris, this volume represents the culmination of more than a decade of work dedicated to uncovering the practicalities of language evolution in a social setting. The book is divided into three sections. An introductory section comprises a Foreword and an Introduction, both by Steels, that set out the direction and the theoretical framework for the remaining papers. Part 1 describes experiments in vocabulary evolution and Part 2 details how grammatical features evolve in experiments in the same framework. Each experiment enhances results gained in previous experiments.

The Foreword places the volume in its historical context by stressing that the question of language evolution is almost as plagued by speculation today as in 1995, when Steels launched this research project. Because there is no fossil record and because we cannot allow any modern language to represent languages as they first emerged, we can only be guided by general principles of evolution when theorising the evolution of languages. Steels and his team have since synthesised an approach to language evolution that attempts to simulate the evolution of language in a cultural context by using computational agents, typically embodied as robots. The Foreword also summarises each chapter.

Chapter 1, “Self-organization and selection in cultural language evolution” by Steels, outlines the theoretical framework for the empirical descriptions in the remaining chapters. Steels demands that any theoretical description of language evolution be biologically feasible, demonstrate advantage to social reproduction, and adapt to cultural change. Language in this model is assumed to be open-ended, distributed, and transmitted non-telepathically. The key aspects of an evolutionary theory that are applied to language are fundamentally functional, i.e., Does language succeed in communicating? Agents apply general strategies that adapt language for optimum expressive adequacy, cognitive effort, learnability and social conformity. The repeated application of these strategies to instances of communicative events produces a language system based on the probability of communicative success. The language system is the combination of the general cognitive capabilities of routine processing and meta-analysis. Ready-made responses may be available to a speaker, but analysis is required to evaluate those responses. Where self-evaluation indicates a lack of success, a repair is introduced. Repair actions may require a reframing of the chosen sentence, the selection of an alternative lexical item, or the creation of a new item or structure. Self-evaluation is possible because of a routine termed ‘re-entry’ (i.e. a process that matches the mirror-neuron hypothesis; see Rizzolatti and Craighero, 2004), which allows the speaker to practice the communicative effect of the chosen sentence before it is articulated by acting as the hearer in an internal process.

While the language system adapts, constrained by language strategies, language items emerge through a self-enforcing cumulative process of invention, trial, and alignment between agents. As with repair, alignment is central to the self-organising character of language. Alignment is the social enaction of frequency, such that the communicatively successful use of a language item increases the likelihood of it being adopted by other agents. This process is demonstrated throughout the volume in various experiments. To further strengthen the centrality of self-organisation in language evolution, Steels also employs the principle of 'structural coupling' (Maturana, 2002), which facilitates alignment through linguistic transmission due to the structure of an organism and its interaction with the living, non-living and linguistic environment, without the need for intention or a central authority.

The key issue for Steels is to provide empirical evidence for the theoretical framework sketched here. Contemporary evolutionary linguistic processes, such as creolisation, can shed light on how language evolves, as can placing linguistically-competent subjects into a context where new language must be invented to complete a communicative task. However, Steels and his collaborators choose to model evolutionary processes computationally and robotically, using embodied agents to enact language games. Throughout the volume, robots engage in: acquisition experiments, where one linguistically competent robot passes on a linguistic system to another robot with a pidgin version of the language, through tutoring, although neither robot knows which has the full version; emergent experiments, where both robot agents, using the strategies described above, collaborate to converge on a non-predetermined stable linguistic system; and reconstruction experiments, where strategies are varied by agents to simulate known linguistic evolution. The remaining chapters describe these experiments for selected vocabulary (Part 1) and grammatical (Part 2) features.

Steels and Martin Loetzsch start Part 1 with the simplest language game: the naming game. In the “non-grounded” version of this emergent experiment, two agents share the same viewpoint of a set of objects. The speaker offers a name for an object, to which the hearer points. If the hearer matches the speaker’s object, a new round is played. However, a number of repairs may occur. The speaker may identify an object with no known name, in which case it has to invent one. The hearer may not know the word, so it guesses the object. If the guess is correct, the new word is remembered, but if it is incorrect, the speaker points out the object, and the new word is remembered. If the hearer knows a different word for the same object, scores are given to the different words so that, through usage, agents converge on agreed words. Thus, in one experiment, twelve words for five objects after 50 games become five to six words, on average, after 200 games. In the “grounded” version of the game, the agents are mobile and may see the same objects from different angles. Identifying objects through luminance, yellow/blue and red/green scores, x and y coordinates, and height and width measurements, agents store prototypes of objects which they then collaborate to name with other agents, using similar strategies and repairs as in the previous game. Aggregate results produce close to 100% communicative success after 1,000 games producing 20 terms after 18 views of 10 objects. Adding the ability to both track moved objects and update prototype models results in about 90% success with 11 terms from 1,500 games after 16 views of 10 objects. That is, these two learning heuristics produce far less ambiguity and synonymy.

In “Language Strategies for Color”, Joris Bleys engages robot agents in naming games for colour, thereby accounting for how categories emerge from a natural spectrum. Agents carry out the same language games as in the previous experiment, but here the objects are distinguishable only by colour. Robot agents use a learning strategy that adjusts, rather than replaces, the current prototypical colour towards the speaker’s use of the colour word whenever communication is unsuccessful. Using English words based on scores for brightness, red/green and yellow/blue scores, robot agents score about 83% communicative success, matching baseline or target scores set by human agents. In an emergent experiment using only hue (or brightness), robot agents achieve about 72% success. To make the experiments more closely match natural language, Bleys also investigates graded membership of colour categories (e.g. “only slightly”, “somewhat” or “very” red). In a reconstruction experiment, robot agents produce words that were “qualitatively similar” (p.74) to their human counterparts in baseline data. Similarly, in acquisition experiments, robot agents demonstrate communicative success at rates marginally below humans. An emergent experiment for colour produces almost 95% communicative success with little variance for 5 words after about 15,00 games. The impressive results for graded membership demonstrate another important aspect of these evolutionary experiments: language strategies adapt to give selective advantage. In this case, graded membership of colours allows a higher rate of success than brightness-only or hue-and-brightness systems.

The experiments in the next two chapters, “Emergent mirror systems for body language” by Steels and Michael Spranger and “The co-evolution of basic spatial terms and categories” by Spranger, add complexity to the linguistic models developed in the previous two chapters by adding verbal and adverbial options (Steel and Spranger) and prepositional meanings (Spranger). Spranger’s experimental embodied-robotic subjects achieve 98% communicative success when reconstructing German spatial terms. Steel and Spranger claim that “It is only by the full integration of all aspects of language with sophisticated sensory-motor intelligence that agents were able to arrive at a shared communicative system that is adequate for the game” (107) of correctly ordering a fellow robot agent to strike a particular pose. That is, communicative success is achieved by: grounding the agents in a sensory experience relative to their own body and its parts; employing a prototypical, rather than categorical, approach to language; simulating mirror neurons (by enabling robots to simulate and monitor, without enacting, a motor programme); and providing feedback loops for the motor system.

Part 1 culminates in the chapter “Multi-dimensional meanings in lexical formation”, by Pieter Wellens and Loetzsch, which attempts to simulate a more natural environment for lexical emergence and demonstrate the adaptive benefits of the strategies adopted in the studies in this volume. The language games played by robot agents in the preceding chapters all focus on one aspect of language, but this does not reflect natural language use, when speakers must select the most suitable linguistic features to distinguish objects. The most favourable results are obtained when agents use a probability-based ‘Adaptive Strategy’ for word learning, whereby a fuzzy-logic algorithm for ‘best fit’ is used in naming objects as speaker or hearer. In experiments where 25 agents able to distinguish 16 features per object play 4,000 games each, totalling 50,000 games over 10 repetitions, the agents achieve 90% communicative success after 10,000 games, and approach a 98% success rate after 30,000 games. Another measure, lexicon coherence, which quantifies the alignment between agents’ lexicons at any time, reaches 0.4 after 10,000 games and averages only as high as 0.45 on a scale of -1 to +1 after 50,000 games. This reflects natural language, where high levels of communicative success are achieved even when agents do not totally agree on word meanings.

Part Two of the book, ‘Emergence of Grammatical Systems’, opens with Remi van Trijp’s ‘The evolution of case systems for marking event structure’, which posits three bold hypotheses: 1. “Case evolves because it has selective advantage for communication” (170); 2. case emerges when a population shares a ‘case strategy’; and 3. “Case markers can be repurposed for a different language system if the original selective advantages of a case system have been ‘usurped’ by more dominant, competing systems in the language” (170). In experiments where one robot agent describes a scene that the two agents have just watched together, robot agents acquire the case system for German, although van Trijp rejects the need for 'a priori' grammatical categories. After 5,000 games, coherence scores are above 0.95, the language system is highly systematic, and cognitive effort is at a minimum, thus providing support for the first hypothesis. Moreover, the evolution of the Spanish personal pronoun system is reconstructed in experiments that provide evidence for hypotheses 2 and 3 above. As with native speakers, grammatical variation is accommodated by robot agents who produce language with preferences for certain structures. Similarly, subsequent experiments demonstrate a paradigm shift in the population, with preferences moving from one system to another. In the conclusion, van Trijp is careful to emphasise that these experiments demonstrate a high level of communicative success using general shared cognitive strategies – typically, “analogical reasoning or similarity-based categorization” (202).

In “Emergent functional grammar for space”, Spranger and Steels demonstrate the selective advantage of grammaticalising spatial relationships over the solely lexical variant in experiments that reconstruct German and that self-organise into an emergent system. Crucially, they show how a semantically-oriented strategy towards grammaticalised spatial relationships requires less cognitive effort for greater communicative success. Similarly, Katrien Beuls, Steels and Sebastian Höfer’s experiments into “The emergence of internal agreement systems” produce results that reduce cognitive effort and ambiguity by grouping related words into groups or phrases. Kateryna Gerasymova, Spranger and Beuls investigate the Russian system of Aktionsarten in “A language strategy for aspect”. Although the Russian system of aspect is considered complex and elaborate, robot agents are able to reconstruct and acquire the system, partially aided by the ability to accept holophrases (a learned combination of words) for later analysis. Robot agents then demonstrate how an entirely new aspect system can emerge. As in the experiments by Wellens and Loetzsch, the final chapter ''The emergence of quantifiers'', by Simon Pauw and Joseph Hilferty, demonstrates the selective advantage of fuzzy categories by focusing on quantifying expressions. Experiments in acquisition and formation compare the alternative strategies of absolute quantification and scalable quantification, resulting in the conclusion that the more unpredictable the environment, the more likely a scalable strategy will prevail.

EVALUATION

Although each paper has different authors, the volume exhibits both a remarkable sense of consistency and a clear sense of progression from one chapter to the next. The research reveals a sense of direction shared by Steels and the other contributors that is laid out in Chapter One. In fact, it is advisable to read Chapter One again after examining the results of later experiments, in order to fully appreciate the significance of the bold approach taken by this team of researchers.

The greatest danger of depending on functional explanations to support a hypothesis is that evidence can only be interpreted as supporting an inert status quo. Fortunately, Steels and colleagues avoid this theoretical blind alley by incorporating the dynamics of alignment and the explanations and mechanisms for linguistic change. For instance, in van Trijp’s chapter, experimental evidence provides support for the hypothesis that the advantages provided by grammatical case in Spanish have been replaced by other grammatical features, freeing case markers to function in new ways. Perhaps my only concern with some of the papers in the volume is that there is an over-reliance on formal, rather than functional models of language. While some functional models may be difficult to model computationally, there are solutions, such as Halliday and Matthiessen (1999), which may provide the research team with grammatical models more aligned with the non-representational approach to language that is central to the research reported here.

This book and other experiments by the same team provide empirical evidence for the emergence of language based on evolutionary principles, on what we currently understand about brain structure and organisation (e.g. Edelman 1999; 2004) and, significantly, without the need for language-specific acquisition strategies; in all of the experiments here, the learning strategies employed are general cognitive strategies rather than language-specific. The experiments repeatedly demonstrate that: language can emerge without a priori conditions; current language systems can be aligned within a community through structural coupling; known developments in language can be modelled in embodied robotic agents with simulated mirror neurons; and language functions probabilistically, not categorically. I am unaware of any other series of falsifiable experiments that provide verifiable evidence to counter these conclusions, despite many theoretical claims to the contrary. Consequently, this volume should be of value to anyone interested in language evolution, in the application of natural languages to robotic agents, and in general linguistic theory.

REFERENCES

Edelman, G.M. 1999. Building a picture of the brain. Annals of the New York Academy of Sciences 882 June 1999, p.68-89

Edelman, G.M. 2004. Wider Than the Sky - The Phenomenal Gift of Consciousness. New Haven: Yale University Press

Halliday, M.A.K. and Matthiessen, C.M.I.M. 1999. Construing Experience through Meaning: A Language-based Approach to Cognition. London: Continuum

Maturana Romesin, H. 2002. Autopoiesis, Structural coupling and cognition: A history of these and other notions in the biology of cognition. Cybernetics and Human Knowing 9(4), pp.5-34

Rizzolatti, G. and Craighero, L. 2004. The Mirror-Neuron system. Annual Review of Neuroscience 27, pp.169-92

Nick Moore has worked in Brazil, Oman, Turkey, the UAE and the UK with students and teachers of English as a foreign language, English for specific and academic purposes, and linguistics. His PhD in applied linguistics from the University of Liverpool addressed information structure in written English. His other research interests include systemic functional linguistics, corpus linguistics, theories of embodiment, lexis and skills in language teaching, and reading programs. He is the co-editor of 'READ', maintains a blog on language, linguistics and learning at najmoore.blogspot.com and has recently joined the TESOL unit at Sheffield Hallam University.


The review for this book is posted here on linguistlist.org. 

Monday, May 7, 2012

Everett Update

Dan Everett's been busy of late. He has been involved in a documentary about his beloved Pirahã, and has a new book out. The documentary is called "The Grammar of Happiness" and the book is called "Language The Cultural Tool." Naturally, all this activity is firmly directed against the school of Chomsky and against Pinker's multi-million best-sellers. Everett appears intent intent on publicly discrediting the generativist/minimalist school. 
Trailer for "The Grammar of Happiness"

The debate rages on about whether the Pirahã language has recursion or not (and if it does not, does it really matter), but just to stoke the fires  higher, Everett has published a new book, called "Language The Cultural Tool." Yes, that's right. You would be hard pushed to pick a phrase that more succinctly says "No. Syntax is not autonomous." For years, generativists have been misrepresenting the ideas of Whorf and Sapir, not least in combining them into a mythical Whorf-Sapir hypothesis that simply does not exist. I hope that Everett is able to bring the debate back onto neutral ground and really tackle the question of how our language construes our perception of the world. For all his image as a 'radical' or 'the U.S. dissident', Chomsky's linguistics is deeply ideologically conservative (Chris Knight explains this very well in Weekly Worker 655, 656 and 657). Maintaining that you do not need to analyse language linguistically in order to identify its power structures, or that habitual language use does not blind one to the legitimacy of incumbent power structures, contributes to the obscuring of the ideological role of language.

You can find reviews of the new book from New York Times and The Guardian, among others.

On a separate but related note, I was on DubaiEye's 'Talking of Books' programme on June 9th, where I  championed Everett's earlier popular book "Don't Sleep there are Snakes" (reviewed in an earlier blog). You can hear most of that segment of the show on Grooveshark.

Thursday, May 3, 2012

Bring Hope to the Bonobos - Again

The Great Ape Trust desperately needs your help to keep their research and the apes alive. Donate to bonobohope.org if you love language, animals or Des Moines, Iowa. Hey, Bill Bryson, I think that must mean you. Anybody have his number?? Join me, Bill Greaves and Peter Gabriel in trying to keep Sue Savage-Rumbaugh's great work going. We do not want another Nim!


More Video links: 
BBC: Super Smart Animals (Great Ape Trust segment starts at time 50:20) 
Oprah Show: Kanzi the talking Ape 
Anderson Cooper (CNN): Anderson as the Easter Bunny (with Kanzi) 
60 Minutes (in Australia): Talk to the Animals
And the latest appeal from Sue & The trust:

Update: I am very happy to report that this year's Target has been achieved. 

Monday, April 9, 2012

Fry's Planet Word on DVD - A Brief Review

As with any 'popularisation' of a subject, academics can easily take a swipe at mass media explanations of their subject. In this review I will try to avoid taking cheap shots (unless the temptation is too great) and attempt to keep an open mind on how well Fry has done in representing the subject of linguistics to the general public as, I believe, this was his intention. I will review each episode and then finish with an overview.

Episode 1 - Babel
...And he's off: in no time at all, we are sent from Stephen Fry's comfy documentary-world study to Kenya to meet the Turkana, back to the London suburbs to meet a typical toddler, Ruby, and off to Leipzig (twice). Before you can say hello in 25 languages, we are already pondering a wide range of linguistic dilemmas. We also catch glimpses of Nim Chimpsky and the famous chattering YouTube twins. To help us out, we visit a range of experts. If I could have anyone in the world to talk about state of the art theories of language development, I would have one person at the top of my wishlist: Michael Tomasello. He adds much-needed balance both to the academic study of language development and the programme itself. Then, very quickly we are back on the slippery search for the language gene, aka FOXP2, and only just understand that this really cannot be the whole answer.
While we do not learn very much about Stephen Fry's brain scans in an fMRI, we do learn that the people he chose from UCL have a very balanced, realistic view of what they are able to achieve with these tools. I suppose if you are going to prepare a documentary on language, you have to include Stephen Pinker, if only because more Joe Publics have read his books than any others on language. Thankfully, Pinker does not get it all his own way. At the end of the episode, we have been given a fairly good overall picture of language development and been introduced to  issues of language versus animal communication, language proliferation, decay and death, and the long, long way we still have to go to even start to understand language. All the time, no matter what you may think of the presenter, Fry clearly enjoys language and relishes the challenge of circumscribing the subject. As an introduction of language study to the completely uninitiated, this is a good start.


Episode 2 - Identity
This episode deals with the typically sociolinguistic topics of accent, language decay and identity.
We start with an investigation of the myriad accents of Yorkshire, guided by a poet from Barnsley, we give Fry a few moments to exhibit his control of accents on an 'accent forecast of the UK' made to resemble a BBC weather forecast, and then we land in Newcastle, where we hear 'chirpy' Geordie call centre operatives and their PR manager. Then we're whisked over to Connemara where we hear some Irish (or Irish Gaelic if you prefer) and find out how the young and older feel about their language which was brought back from the brink of extinction. At this point we do touch on the serious issue of language decay, identity and "linguicide", with a cheap swipe at L'Académie française for being so imperial for so long. We also look at the re-birth of Hebrew, where we at last meet a linguist (only the 2nd in this programme) whose thesis is that Hebrew still retains large parts of Yiddish. (I do not know if Fry is Jewish, but in this episode he goes out of his way to be nice to them in London, New York and Jerusalem.) Finally we compare how Irish, Breton, Basque, Hebrew, Oc and Turkana resist the threat of Globish (that's global English).
One other point of interest in the programme is the debate around how far your culture affects your language, and vice versa, with Stanford Russian linguist Lera Boroditsky discussing how masculine and feminine nouns in gendered languages affect the way that some speakers describe objects that carry different genders in different languages. Fry later admits to supporting the Chomskian line, that all languages are ultimately similar and so, after allowing such a poor misrepresentation, does a double disservice to the so-called Sapir-Whorif hypothesis.
After the Frying start of episode 1, episode 2 is very disappointing. I do not think that this is due to my personal lack of interest in the issue of Identity (which I think covers a multitude of academic sins), but because the head count of experts - famous or otherwise - is much lower in this episode, and Fry's inexhaustible enthusiasm is an insufficient replacement for real facts.


Episode 3 - Uses and Abuses
The primary aim of this episode appears to be to cram in as many words that are normally banned on the BBC as possible. It does quite a good job with copious fucks, plenty of bollocks and a smattering of cunts. In terms of academic head counts, this episode does a much better job than number 2, except Stephen Pinker crops up a number of times spouting off on subjects that he really has no expertise in - a role in which he has become quite an expert!
Although  Fry's approach could easily be dismissed by people working in the fields of sociolinguistics and humor studies (which he refers to as rather humorless), we must never forget that this is a television programme.
An experiment involving the actor Brian Blessed and a large tub of icy water is particularly unscientific, but it makes good television and is slightly related to more serious research. Admittedly this episode does descend into a promo-video for Fry's favourite issues (Judaism, homosexuality, racism etc.), but generally it also makes a point; although Fry & co. smatter their speech with expletives, none of them can bring themselves to say nigger in any form other than "the n-word." So, even for people who can cuss and swear willy-nilly, there are still some taboo words. Stephen K. Amos manages, just, but explains that it still retains an insulting meaning for him due to personal experiences. Hardly scientific. An improvement on episode 2, but still not as accurate or rigorous as episode 1.


Episode 4 - Spreading the Word
Episode 4 is all about writing, and is of a much higher standard than the previous 2 episodes. Although I would not agree with all that we see in this episode, I would say that Fry and his team have done a better job of researching the key issues in writing. We have a wide range of suitable experts from typesetters in Norwich to the inventor of Pinyin in Beijing.
When Fry wants to learn about the origins of writing he finds an expert in cuneiform in the British museum, who shows him how it is done, and even poses in front of THE Rosetta Stone (best not to ask why it's in the British Museum, though!!). He is back in Jerusalem to look at, be told off for touching, and witness digital imaging techniques for restoring the "Dead Sea Scrolls."
He also manages to trace the connection between printing, the Age of Reason and wikipedia - not bad for a TV show! He visits the Bodleian library at Oxford University to examine how they are keeping abreast of the digital age and in Harvard's library meets someone who points out that new media do not need to replace previous formats - that the iPad / Kindle etc. are not likely to replace the book, but both will develop alongside each other just as radio did not replace newspapers and was not replaced by TV.
As with other episodes Fry adds his own pet theories, likes and dislikes but he also places developments such as printing in a social context, providing a good balance of enthusiasm and restraint on a subject that easily leads to hyperbole. I also support his call to support the libraries of the U.K. and the world, no matter what formats are being preserved - buildings dedicated to the pursuit of knowledge through reading are the bedrock of civilizations!


Episode 5 - The Power and The Glory
So in the ultimate episode we learn that the ultimate purpose of language is...
literature. 
Fry explains why he loves a range of writers, from Joyce to Wodehouse to Orwell, heaping the greatest praise on Shakespeare - he even manages to find a French actor to admit that he would rather play Shakespeare than any of the lesser French playwrights. Fry just could not resist one last jab at the French before the series finishes. Some old friends are back, such as Brian Blessed and the Turkana villagers, as well as some new faces, including David Tenant giving us some of his take on Hamlet.
The episode is just as busy and full of locations, interviews and Fry's opinions as the others, but offers no new information on linguistics or even language studies. For this reason it is the most disappointing episode - at least episode 2 was related to aspects of sociolinguistics and the hot topic of identity. All of the science disappears and we left with the absolute relativism of everyone's opinion is just as good as each other, which is clearly not the case. Just ask a Cambridge Don!


Episodes 1-5 - Planet Word

All in all, Planet Word is very uneven. It manages to combine wit and fact, controversy and error. At times, it is highly perceptive and at others completely misleading. To be fair, this is TV. It is not intended to be lectures 1-5 in a course in linguistics. Evaluating the series from an academic perspective is completely unfair. Whatever Fry offers, it must work well on the screen - hence the frequent scene changes (often for no reason), cuts to the fake study for a talk-to-camera and a heavy dependence on interviews with experts (used in the loosest sense where Stephen Pinker is concerned).
What we need most from this series, perhaps, is for the general public to gain some understanding of language studies or even be inspired to look further into the subject, especially if they are young and are considering what to study at university. I believe Fry has succeeded to some extent in providing a TV series that engages with its audience, entertains and informs. A wide range of linguistic issues, perspectives and facts are offered with a minimum of effort on the part of the viewer - no mean feat. Only the last episode could be considered misleading. Certainly I do not agree with a lot of what he claims throughout the series, but this is his show not mine!
Find yourself a copy of the DVD, or even pick up the book, and see if you could find better ways to make linguistics appeal to more people who have never considered studying language before. It will not improve the programme but it may just help you, if you are a lover of language, to explain your interest to others. Spread the word - Planet Word.

Sunday, March 11, 2012

Iain McGilchrist - The Divided Brain

This video provides an up-to-date discussion of what happens in the hemispheres of the brain. It quickly rejects the old emotional-rational, language-mathematical & other divisions, but explains just what happens in the two halves, and why.
(If the video fails, use this link)
This is another RSAanimate video produced by the excellent Cognitive Media team, who also produce A0 size digital files and prints of the final version of the talk. There are many other similar videos, but for content and visuals this is my favourite so far. You can also download an iphone* or android app from RSA that grants you access to the animated talks, the original lectures of these and more talks, and much more.

* Not that you would know what this post is about on your iphone because it does not support flash video.

Monday, March 5, 2012

Clark & Lappin - Linguistic Nativism and the Poverty of the Stimulus

Another post (back in November) notified the publication of "Linguistic Nativism and the Poverty of Stimulus" by Alex Clark & Shalom Lappin. This post is a review of the book. Below is the extended version, while this link sends you to the edited version on linguistlist.


AUTHORS: Clark, Alexander and Lappin, Shalom
TITLE: Linguistic Nativism and the Poverty of the Stimulus 
PUBLISHER: Wiley-Blackwell
YEAR: 2011

Nick Moore, Khalifa University, United Arab Emirates

In 12 chapters, with front material, contents, a preface, reference, an author index and a subject index, Alexander Clark and Shlalom Lappin’s “Linguisic Nativism and the Poverty of the Stimulus” tackles a key issue in linguistics over 260 pages. The book is intended for a general linguistics audience, but the reader needs some familiarity with basic concepts in formal linguistics, at least an elementary understanding of computational linguistics, and enough statistical, programming or mathematical knowledge not to shy away from algorithms. It is written, however, for a wide range of undergraduate, graduate and practicing linguists, particularly researchers working in formal grammar, computational linguistics and linguistic theory.

The main aim of the book is to replace the view that humans have an innate bias towards learning language that is specific to language with the view that the innate bias towards language acquisition depends on abilities that are used in other domains of learning. The first view is characterised as the argument for a strong bias, or linguistic nativism, while the second view is characterised as a weak bias or domain-general view. The principle line of argument is that computational, statistical and machine-learning methods demonstrate superior success in modeling, describing and explaining language acquisition, especially when compared to studies from formal linguistic models based on strong bias arguments, typically from Universal Grammar, Principles and Parameters, the Minimalist Program and other theories inspired by the work of Noam Chomsky.


Summary

Chapter 1, the introduction, establishes the boundaries of the discussion for the book. The authors focus throughout on a viable computationally-explicit model of language acquisition. While briefly presenting arguments on the evolutionary and biological nature of linguistic nativism, they rarely consider these questions again in this book. Their first of many criticisms of Universal Grammar (UG) theories proposed and inspired by Noam Chomsky is the Minimalist Program appears to disregard the puzzle of acquisition, unlike earlier versions of the theory which placed innate language-specific learning at the core of the theory, in order to explain how children consistently learn language from apparently inadequate data. Clark and Lappin (hereafter C&L) do not dismiss nativism entirely. They point out that it is fairly uncontroversial, first, that humans alone acquire language, and, second, that the environment plays a significant role in determining the language and the level of acquisition. What they intend to establish in the book, however, is that any innate cognitive faculties employed in the acquisition of language are not specific to language, as suggested by Chomsky and the Universal Grammar (UG) framework, but are general cognitive abilities that are employed in other learning tasks. It is from this ‘weak bias’ angle that they critique the Argument from the Poverty of Stimulus (APS).

Chapter 2 focuses on the Argument from the Poverty of Stimulus (APS). The APS is considered central to the nativist position because it provides an explanation for Chomsky’s core assumptions for UG: (1) grammar is rich and abstract; (2) data-driven learning is inadequate; (3) the linguistic data that the child is exposed to is qualitatively and quantitatively degenerate; and (4) the acquired grammar for a given language is uniform despite variations in intelligence. Most of these assumptions are dealt with throughout the book, and many replaced with assumptions from a machine-learning, ‘weak bias’ perspective, although little evidence is supplied to counter assumption three; rather, the reader is referred to other sources. Some APS theories have attacked connectionist approaches to the language acquisition puzzle by proposing that frequency alone would predict a radically different learning order than that observed. C&L counter that an argument against a connectionist claim does not prove the UG assumption of linguistic nativism correct – the strong bias towards language-specific learning mechanisms remains unproven – and they contend that the puzzle of language acquisition can be more explicitly delineated by computational methods than the proposals so far provided by the various UG frameworks. Reviewing the UG evidence for a strong bias, C&L claim that many arguments become self-serving: “a particular grammatical representation is not motivated by the APS, but rather it becomes an argument for the APS.” (p.33) Two examples of the APS – auxiliary inversion and pronominal ‘one’ – are provided as cases in point. In this chapter C&L introduce machine-learning alternatives to linguistic nativism based on distributional, or relative frequency, criteria and a Bayesian statistical model to account for these same learning phenomenon. These alternatives are elucidated further in the book.

Chapter 3 examines the extent to which the stimulus (accurately, the Primary Linguistic Data) really is impoverished. C&L avoid arguments between the more empirically-demonstrated richer linguistic environment and the naively-assumed impoverished input (the reader is referred to MacWhinney and others), but examine a key question for the computational modelling of the learning process: the existence or prevalence of negative evidence in the learning process. Even a small amount of negative evidence, for instance through reformulation, can make a significant difference to the problem of learning, and C&L allow for far less negative evidence than has been demonstrated in a range of corpora of child-directed speech. Other indirect negative evidence, such as the non-existence of hypothesised structures, can also significantly alter the learning task assuming that the learner is free to make these hypotheses. C&L challenge the premise of no negative evidence, partly because it forms such a central tenet of Gold’s “Identification In the Limit” – a theory of language acquisition that provides considerable support for the UG position of linguistic nativism. Gold’s highly-influential study argues that because learning cannot succeed within the cognitive, developmental and time limits imposed, then children must have prior knowledge of the language system. For instance, Gold claims that since the Primary Linguistic Data does not contain structural information, such as tagging for correctness, the knowledge that children rapidly acquire (e.g. knowing if a string is grammatical) can only be explained by strong linguistic nativism. However, C&L point out that this view is partly a consequence of ignoring non-syntactic information in the learning environment. Maybe it is not possible to explain structural acquisition independently, but the addition of the semantic, pragmatic and prosodic information of the learning context cannot be ignored in the learning process without producing a circular argument. The “Identification In the Limit” theory is examined further in the following chapter.

Clark and Lappin’s overall goal is to establish formal, computational descriptions of the language learning process that can be demonstrated to be viable and feasible, although they are keen to point out that demonstrating the tractability of the learning problem does not equate to modelling the actual neurological or psychological processes that may be employed in language acquisition. Chapter 4 discusses a major argument against the machine learning approach: Gold’s “Identification In the Limit” theory, which concludes that language acquisition is not viable for a ‘learning machine’ and so only linguistic nativism can explain success in language acquisition. C&L reject a number of assumptions in Gold’s model. They do not agree that presentation of the target language should be considered unlimited, as this allows for the unnatural condition of the malevolent presentation of data – intentionally withholding crucial evidence samples and offering evidence in a sequence detrimental to learning. They reject Gold’s lack of time limitation placed on learning. They reject the impossibility of learners querying the grammaticality of a string. They reject the argument that ‘side’ information – information relating to the pragmatics and semantics of the learning context – has no influence on learning syntax, and they provide further evidence to reject the view that learning is through positive evidence only. Perhaps the most significant assumption made in the Gold model that C&L reject is the insistence on limiting the hypothesis space available to the learner. Rather, C&L insist that it is the learnable class of a language that is limited, while the learner is free to form unlimited hypotheses on the limited language. It seems that Gold’s approach is an argument for APS which does not consider alternative approaches: “The argument for the subset principle rests on similar misconceptions as the argument that a target-learnable class must be known to the learner.” (p.97)

Proceeding from a critique of the UG position of a strong bias towards linguistic nativism, C&L begin to introduce their alternative machine-learning approach in chapter 5, “Probabilistic Learning Theory”. The initial step in the weak bias argument is to replace a binary definition of a convergent grammar, typical in UG threories, with a probabilistic definition as this more accurately reflects natural learning processes. C&L on this, and a number of other occasions, object to the simplistic lines of argument employed by Chomsky and his followers in their rejection of statistically-based models of learning. While it may be true that the primitive statistical models critiqued by Chomsky are incapable of producing a satisfactory distinction between grammatical and ungrammatical strings, this does not prove that all statistical methods are inferior to UG descriptions: “the failure of a simple word bigram model to predict a difference in probability between observed events does not entail that statistical language modeling methods in general are unable to handle grammar induction.” (p.101) Consequently C&L introduce a range of statistical methods that they propose can better represent the nature of language acquisition than the under-specified domain-specific mechanisms presumed in UG theories. Central to a plausible probabilistic approach to modelling language is the distinction between simple measures of frequency and distributional measures. Here C&L are proposing that a learner will hypothesise the likelihood of a sequence, based on observations, in order to converge on the most likely grammar. Recent studies using such probability-based grammars are reported in this and later chapters to offer very reliable results, even approaching a reliability factor close to 90%. This general framework uses Probably Approximately Correct (PAC) learning algorithms. PAC algorithms predict efficient, time-limited learning, without requiring the learner to know that their grammar has converged on the correct grammar. However, traditional PAC models have problems, including the requirement that data samples are labelled and the conditions under which a language can be learned, which make them unlikely candidates for reliable models of natural language processing. These issues are dealt with in the following chapters by modifying PAC algorithms in order to better reflect the conditions of natural language processing.

Replacing Gold’s paradigm and the PAC learning with three key assumptions (1. language data is presented to the learner unlabelled; 2. the data includes a proportion of ungrammatical sentences; and 3. efficient learning takes place despite negative examples) allows C&L to introduce the Disjoint Distribution Assumption to more accurately reflect natural language learning, in chapter 6. This probabilistic algorithm depends on the observed distribution of segmented strings, and on the adoption of the principle of converging on a probabilistic grammar (a string is probably correct) in place of a categorical grammar (a string is definitely correct). Using a distributional measure ensures that the probability of any utterance will never be zero, allowing for errors in the presented language data, but each string will be measured against its observed likelihood of distribution. In fact, this model predicts over-generalisation and under-generalisation in the learner’s output because, with an unlimited hypothesis space, “It is the ungrammatical strings that an incorrect hypothesis would wrongly predict to be grammatical, and of high probability, that provide crucial data for learning.” (p.133) The addition of word class distributional data – the likelihood of a certain word class in preceding and succeeding position – also ensures greater reliability of judging the probability of a string being grammatical.

A major aim of this book is to provide a computational account of the language learning puzzle that may not necessarily replicate natural language acquisition, but will make the problem tractable – possible within the defined assumptions. It is the contention of the authors that UG theories have made the wrong assumptions in relation to the learning task and the learning conditions, and in chapter 7 “Computational Complexity and Efficient Learning” they set out the assumptions that allow learning to be efficient without positing a strong bias toward linguistic nativism. To achieve this, they examine the amount and complexity of the data, and the nature and constraints of the learning task, in order to propose learning algorithms that can simulate learning under these conditions, while warning that the purpose here is to demonstrate the possibility in a computational environment not the actual psychological processes that enable human language learning. That is, demonstrating a computational or a linguistic model of learning and language does not entail its psychological reality. A central assumption that is essential to C&L’s machine learning approach is that the input data is not homogenous, resulting in some parts of language being ‘more learnable’ than others. Using a standard domain-general capacity to cluster, the language learner can focus on the easier learning tasks leaving the more difficult parts of grammar for later. More complex learning tasks can then be attacked class-by-class according to a Fixed Parameter Tractability algorithm. Ultimately, C&L argue that complex grammatical problems are no better solved by a UG Principles and Parameters approach; the learning problem remains just as complex and learning need not be achieved any more efficiently. Thus, when UG theories use ‘strong bias’ position as the only argument to deal with complexity, they have not solved the problem posed by a seemingly intractable learning task.

If we are to reject the presumption of the strong bias in linguistic nativism, we need to be confident that its replacement can produce reliable results. Chapter 8 starts to provide those results, illustrated in a range of proposed algorithms. The process of hypothesis generation in Gold’s ‘Identification In the Limit’, a key support of UG, is described as being close to random, and consequently “hopelessly inefficient” (p.153). Various replacements that have been tested, initially in non-linguistic contexts, include (Non-)Deterministic Finite State Automata. These algorithms have then proved effective in restricted language learning contexts. Simple distributional and statistical (including hidden Markov models) learning algorithms offer promising results, but must be adapted to also simulate grammar deduction. Lattice based formalisms are offered as one prospect to “demonstrate tractable learnability for a nontrivial subclass of context-sensitive representations.” (p.161)

Despite promising results, there are still objections to distributional models, and these are countered in chapter 9, “Grammar Induction through Implemented Machine Learning,” which describes the results of real algorithms working on real data. In general, learning algorithms are tested against a ‘gold standard.’ Typically the algorithm performs a parsing, tagging, segmenting or similar task on a corpus, which may or may not be labelled in some way, and the results are measured against a previously-annotated version of the corpus. Corpora in these experiments tend to be samples of English intended for adults – such as the extracts from the Wall Street Journal included in the Penn Treebank (Marcus, Marcinkiewicz and Santorini, 1993). Success is measured by how closely the algorithm matches the previous results and is typically presented as a percentage. Learning algorithms can be divided into “supervised” – requiring the corpus to carry some form of information such as part of speech tagging – and “unsupervised” – working on a ‘bare’ corpus of language.  Not surprisingly, supervised learning algorithms, such as the Maximum Likelihood Estimate, match the ‘gold standard’ in about 88-91% of cases. More surprising, perhaps, are the success rates of unsupervised learning algorithms in word segmentation, in learning word classes and morphology, and in parsing. For instance, parsing algorithms match from 52 to 87% of cases when compared to the ‘gold standard’ of previously-annotated corpora. However, close examination of results often reveals explainable differences – that is, the algorithm produces results which are plausible even if they do not match the human annotation. In brief, various experiments have demonstrated that in the vast majority of cases, learning algorithms can accurately segment and categorise adult language without guidance, and make suitable hypotheses about their grammatical role.

Chapter 10 returns to UG models of language learning, and ‘Principles and Parameters’ arguments in particular, in order to compare them with statistical models of learning. C&L claim that the strong bias in this UG model would require an innate phonological segmenter and part of speech tagger, and that by limiting the hypothesis and parameter space available to the learner, the language learning task actually becomes far more complex, particularly as the highly abstract nature of UG parameters appear to have very little direct relationship to the primary language data. Tellingly, C&L lament the paucity of theoretical and experimental evidence for the Principles and Parameters (P&P) framework to provide examples of language variation and learning that fit with the proposal that learning one ‘parameter’ in a language automatically entails the learning of a set of features:
“The burden of argument falls on the advocates of the P&P view to construct a workable and empirically adequate type hierarchy that captures the main patterns of language variation. The fact that such a theory has not yet emerged, even in general outline, after so many years of work within the P&P framework provides good grounds for questioning the concept of UG that it presupposes.” (p.184)
Even more worrying for UG supporters of a strong bias towards innate language acquisition is the near-indifference to questions of acquisition in Minimalist Program research, the latest theory in UG. A strong bias towards linguistic nativism requires the human brain to evolve these biases in order to learn language efficiently. However, this places language prior to humans, as part of the environment to which humans must adapt. Christiansen and Chater (2008) are credited by C&L with discrediting this notion: it is language that must adapt to fit the human mind. Seen in this light, it is far more likely that languages adapt to learning biases that have already evolved in the human brain, providing another ‘weak bias’ argument, and so general tendencies in human languages are not examples of Universal Grammar but “emergent properties of the processes through which natural language is adapted over generational cycles of transmission.” (p.185) In place of the UG models, C&L again offer sophisticated statistical models of language learning and generation. Promising results using “Probabilistic Context-Free Grammars” and “Hierarchical Bayesian Models” currently provide the best alternatives to UG models by accounting for language acquisition through comprehensive descriptions and high levels of success in simulations.

In chapter 11, “A Brief Look at Some Biological and Psychological Evidence” C&L quickly review accounts of language learning that support a weak bias. Even where evidence from genetic, biological or psychological studies have been used to support a strong bias, C&L are able to show that this evidence does not necessarily favour nativist arguments. For instance, the FOXP2 gene has been attributed, by Pinker and Jackendoff among others, as the gene responsible for the capacity for human language. In a family where this gene is mutated, all family members suffer a similar severe language disability. C&L point out, though, that this gene is not unique to humans and, as it is responsible for a disruption in motor learning and development in other animals, it is far more likely that this genetic mutation results in a more general learning disability that also severely affects speech and language development. Similarly, psychological disabilities such as Williams Syndrome that may appear to be language-specific are, on closer inspection, related to a range of learning and developmental abnormalities.

In the final chapter, ‘Conclusion,’ C&L review their evidence against the argument from the poverty of the stimulus. They remind us that they are arguing for a weak innate bias towards learning language based on general-domain learning abilities rather than language-specific abilities. They point out that the UG framework has produced few concrete experiments or cases that explain the language variation or acquisition described in theoretical accounts or that “produce an explicit, computationally viable, and psychologically credible account of language acquisition” (p.214). What they have attempted to demonstrate in “Linguistic Nativism and the Poverty of the Stimulus” is that there are explicit computational models of learning that have produced a credible account of learning without requiring language-specific parameters. Although they are far from perfect and much work needs to be done, computational models have already provided a more adequate account than the UG models:
“We hope that if this book establishes only one claim it is to finally expose, as without merit, the claims that Gold’s negative results motivate linguistic nativism. IIL is not an appropriate model. Conditions on learnable classes need not be domain specific. Finally, the learner does not have to (and in general, cannot) restrict its hypothesis to the elements of the learnable class.” (p.215)
Instead C&L advocate the use of domain-general statistical models of language acquisition.


Evaluation
In 12 chapters, Clark and Lappin use “Linguistic Nativisim and the Poverty of Stimulus” to evaluate a key concept in modern linguistics, taking a clearly computational perspective and examining a wide variety of topics. This monograph presupposes familiarity with most of the core issues but does not demand in-depth knowledge of computational linguistics. In such a review I have, naturally, simplified or ignored a number of important arguments presented in this book, and skimmed over significant presentations of learning algorithms. I would suggest however, from this reviewer’s point of view, that C&L present a very cogent and coherent argument.

There are so many sides from which to attack linguistic nativism in general, and the argument from the poverty of stimulus in particular. Opponents have argued that most UG theories are unfalsifiable (e.g. Pullum and Scholz, 2002), that corpora designed to reflect children’s exposure to language demonstrate that the stimulus is not impoverished (e.g. MacWhinney, 2004), that it is absurd to posit the notion that the brain adapted to language, as if language exists in the environment prior to man, rather than language adapting to the general abilities of the brain (Deacon, 1998; Christiansen and Chater, 2008). These arguments, alongside alternatives to linguistic nativism from functional linguistics (e.g. Halliday, 1975; 1993; Halliday and Matthiessen, 1999), are often easily dismissed as being irrelevant to the theory of UG. Some of these arguments are mentioned in this book, but what sets Clark and Lappin’s book apart, and why it must be taken seriously by everyone who proposes some form of UG to explain language acquisition and typology, is that it attacks from within. It claims the very ground claimed by theories of UG. UG attempts to formally and explicitly account for the apparent mismatch between the complexity of the language learning task and the near-universal success of humans in achieving it with such apparently meagre resources. The methods proposed by Clark and Lappin identify what methods could be applied to make the complex task tractable. Specifically, these methods are not restricted to language, but are generally useful learning methods – they are domain-general. If there is one criticism I would make of Clark and Lappin’s argument it is that they do not demonstrate clearly enough – at least to this rather naive reader – just how likely it is that we all use the domain-general learning methods that they propose. For instance, we are left to presume that Probabilistic Approximately Correct and Probabilistic Context-Free Grammar algorithms represent general, non-language specific, models of learning, but this is not demonstrated.

My biggest fear with this approach is that we may be fooled by the apparent sophistication of the tools at our disposal. We need to remember that when using computational tools to help us understand a phenomenon far more complex than computers, we must not allow the tools to force us to see the phenomenon as the tool understands it. It seems more than coincidental that a computational approach to language acquisition mirrors the findings about language that corpus linguistics has revealed; for instance, that language can be viewed as inherently statistically structured. That it can be analysed in this way, or in the form of transformational trees, does not prove that this is how humans learn language. Fortunately, Clark and Lappin are well aware of this trap and frequently warn readers that no matter how well their computational theories may reflect language acquisition facts, the aim of computational models is to demonstrate what is possible, or even likely, not what really happens in the human mind. Until we better understand exactly what neurological processes are actually involved in language acquisition, our task is to try to represent acquisition as best we can. In this endeavour, we have been expertly assisted by Clark and Lappin’s book.

Linguistic Nativism and the Poverty of the Stimulus is a challenging book. It challenges the reader to deal with a range of linguistic, philosophical, mathematical and computational issues. It challenges the reader to remember a dizzying array of acronyms and abbreviations (APS, CFG, DDA, DDL, EFS, GB, HBM, IID, IIL, MP, PCFG, PLD and UG to name but a few). Most of all, it challenges basic concepts in mainstream linguistics. It examines key tenets of UG in the light of advances in machine learning theory and research in the computational modelling of the language acquisition process, and finds them sorely deficient. It exposes so-called proofs supporting the poverty of stimulus, and reveals alternatives that are formally more comprehensive than the explanations previously provided by the linguistic nativism inherent in UG, and empirically more likely to match natural language acquisition processes.


REFERENCES
Christiansen, Morten H. and Chater, Nick. 2008. Language as Shaped by the Brain. Behavioural and Brain Sciences 31. pp.489-558
Deacon, Terrence W. 1998. The Symbolic Species: The Co-Evolution of Language and the Brain. New York: W.W. Norton
Halliday, Michael A.K. 1975. Learning How to Mean: Explorations in the Development of Language. London: Edward Arnold
Halliday, Michael A.K. 1993. Towards a language-based theory of learning. Linguistics and Education 5 p.93-116
Halliday, Michael A.K. & Matthiessen, Christian M.I.M. 1999. Construing Experience Through Meaning. London: Continuum
MacWhinney, Brian. 2004. A Multiple Solution to the Logical Problem of Language Acquisition. Journal of Child Language 31, pp. 883-914
Marcus, Mitchell P., Marcinkiewicz, Mary Ann and Santorini, Beatrice. 1993. Building a Large Annotated Corpus of English: The Penn Treebank. Computational Linguistics19/2, p.313-330
Pullum, Geoffrey K. and Scholz, Barbara C. 2002. Empirical Assessment of Stimulus Poverty Arguments. The Linguistic Review 19, pp.9-50

Saturday, January 28, 2012

Life, the Universe and Everything by Antonio Damasio

Far too often discussions of consciousness are based on speculation and the desire to make presumed facts fit an ideology - yes, I am talking about you Mr. Chomsky. At other times we are fortunate to find people that start with identifiable facts and then attempt to explain them. In this fascinating TED talk Antonio Damasio lets us in on his latest understanding of the human brain, the human mind and the centrality of the self in generating a consciousness. Hold on tight as he covers a very wide range of topics in under 19 minutes, but at the end of the ride I feel much closer to an understanding of what we are.

This talk develops Damasio's earlier theories and is in line with what we know about brains from Edelman, about evolution from Deacon and about language from HallidayThe talk is based loosely on Damasio's latest book "Self Comes to Mind" which is reviewed in Constructivist Foundations 7:1.

Wednesday, November 23, 2011

Descartes' Error

Descartes' Error: Emotion, Reason and the Human Brain
by António R. Damásio
Penguin (Non-Classics), 2005 (first published 1994)

The book starts with neuroscience's cause celebre - a man whose head was pierced by a metal stake that passed through his neck and out of the top of his head. The fact that he survived is astonishing, and is where many neuroscientists in the past have stopped, having proven some point or another about neuroanatomical structure. Damasio not only provides us with gorgeous detail about the tragic accident that resulted in Phineas Gage's custom-made tamping rod exploding through his skull, he also follows Gage after his initial recovery into a tragic story of the downfall of a once-proud man. As well as Gage, Damasio offers many more intriguing stories of brain damage and other ailments that affect the way that we operate in our social environments and in doing so he makes a very strong case for the reintegration into science of emotion. Damasio complains that emotion has been ignored for too long - perhaps because of the over-riding desire to be logical and "scientific".

Even if (as Descartes would have us believe) it is possible to think and act logically, that does not mean that we cannot logically study the emotions. On the contrary, it is the passion, insight, intuition and inspiration that has produced the greatest advances in science - great innovators just knew they were right even when nobody else believed it. More specifically, Damasio argues that it is precisely when people lose their ability to evaluate emotionally that they become paralysed by logical thought. Certain syndromes result in people being unable to choose the right option, even when one may involve losing a job or a friendship. Damasio's answer is to propose a model of thought that gives the emotions a key part in cognitive processes, and demanding that the Dualism so popular in science that follows Descartes is consigned to the history books.

If you are not entirely convinced that emotion plays a part in our most logical thought processes, consider these 2 points:
1. Descartes reasoned that the only truth any of us can be sure of is that "I think therefore I am". One error he made was that he did not take his logical analysis 1 step further. How do we know we think? We feel we know. Without the feeling that we know we would not be so sure that we think. This is not how Descarte's error is explained in the book, but it is what I have learned from it.
2. Our primary sense is not vision. It may be the one we are most aware of using, but vision depends on another sense: Touch. Not touch at the end of the fingertips, but touch as our whole skin. We touch our environment by taking a position within it, and only when we know where we are and how we are situated in our environment can we start using our other senses as comparative measures. We feel who we are and where we are. Without touch we have no awareness, and without awareness we have no thought.

You guessed it... another Goodreads review