Butterfill

The Developing Mind: A Philosophical Introduction

Butterfill, Stephen A. (2020). The Developing Mind: A Philosophical Introduction. https://doi.org/10.4324/9780203758274
Contents
Preface xi
1 Introduction 1
1.1 TwoBreakthroughs....................... 2
1.2 Knowledge............................ 3
1.3 A Crude Picture of the Mind . . . . . . . . . . . . . . . . . . 4
1.4 CoreKnowledge......................... 5
1.5 TwoStories............................ 7
1.6 Development Is Rediscovery . . . . . . . . . . . . . . . . . . 8
I Physical Objects 11
2 Principles of Object Perception 13
2.1 Knowledge of Objects Involves Three Abilities . . . . . . . . 13
2.2 Segmentation .......................... 15
2.3 Principles of Object Perception . . . . . . . . . . . . . . . . . 20
2.4 Conclusion............................ 23
3e Simple View 25
3.1 TheSimpleView......................... 25
3.2 Persistence............................ 27
3.3 Extending the Simple View to Persistence . . . . . . . . . . . 32
3.4 Causal Interactions . . . . . . . . . . . . . . . . . . . . . . . 34
3.5 The Case for the Simple View . . . . . . . . . . . . . . . . . 36
4e Linking Problem 41
4.1 Against the Simple View . . . . . . . . . . . . . . . . . . . . 42
4.2 Further Evidence Against the Simple View . . . . . . . . . . 46
4.3 Things Get Even Worse for the Simple View . . . . . . . . . 48
4.4 The Linking Problem . . . . . . . . . . . . . . . . . . . . . . 50
4.5 Representation Not Knowledge . . . . . . . . . . . . . . . . 51
4.6 Graded Representations? . . . . . . . . . . . . . . . . . . . . 53
v
4.7 Conclusion............................ 55
5 Core Knowledge 57
5.1 What Is Core Knowledge? . . . . . . . . . . . . . . . . . . . 58
5.2 Can Core Knowledge Solve the Linking Problem? . . . . . . 60
5.3 How Not to DeneSomething................. 62
5.4 Will Invoking Modularity Help? . . . . . . . . . . . . . . . . 63
5.5 Conclusion............................ 64
6 Object Indexes and Motor Representations of Objects 67
6.1 Object Indexes in Adult Humans . . . . . . . . . . . . . . . . 68
6.2 Object Indexes and the Principles of Object Perception . . . 70
6.3 The CLSTX Conjecture . . . . . . . . . . . . . . . . . . . . . 74
6.4 SignatureLimits......................... 75
6.5 Knowledge or Core Knowledge or …? . . . . . . . . . . . . . 79
6.6 Against the CLSTX Conjecture . . . . . . . . . . . . . . . . . 80
6.7 Motor Representations of Objects . . . . . . . . . . . . . . . 81
6.8 ConjectureO........................... 83
6.9 Conclusion: Paradox Lost . . . . . . . . . . . . . . . . . . . . 86
7 Metacognitive Feelings 89
7.1 Objection to Conjecture O . . . . . . . . . . . . . . . . . . . 89
7.2 Metacognitive Feelings: A First Example . . . . . . . . . . . 91
7.3 More Metacognitive Feelings . . . . . . . . . . . . . . . . . . 92
7.4 What Is a Metacognitive Feeling? . . . . . . . . . . . . . . . 94
7.5 A Metacognitive Feeling of Surprise? . . . . . . . . . . . . . 96
7.6 Conjecture Om.......................... 97
7.7 Metacognitive Feelings are Intentional Isolators . . . . . . . 99
7.8 Conclusion............................ 101
8 Conclusion to Part I 103
8.1 What Is an Expectation? . . . . . . . . . . . . . . . . . . . . 103
8.2 Core Knowledge: A Lighter Account . . . . . . . . . . . . . 105
8.3 Development Is Rediscovery . . . . . . . . . . . . . . . . . . 106
8.4 How Does Rediscovery Occur? . . . . . . . . . . . . . . . . . 108
9 Innateness 113
9.1 Syntax .............................. 114
9.2 A Poverty of Stimulus Argument . . . . . . . . . . . . . . . . 115
9.3 The Poverty of Poverty of Stimulus Arguments . . . . . . . . 118
9.4 Is Core Knowledge Innate? . . . . . . . . . . . . . . . . . . . 119
9.5 Syntax and Rediscovery . . . . . . . . . . . . . . . . . . . . . 120
9.6 Conclusion............................ 122
vi
Interlude on Innateness 113
II Minds and Actions 125
10 Action 127
10.1 Tracking vs Knowing . . . . . . . . . . . . . . . . . . . . . . 127
10.2 Three-month-olds Track the Goals of Actions . . . . . . . . 128
10.3 Pure Goal Tracking . . . . . . . . . . . . . . . . . . . . . . . 131
10.4 The Teleological Stance . . . . . . . . . . . . . . . . . . . . . 134
10.5 Statistical Regularities . . . . . . . . . . . . . . . . . . . . . . 137
10.6 A Methodological Explanation? . . . . . . . . . . . . . . . . 141
10.7 A Second Puzzle: Acting and Tracking . . . . . . . . . . . . 142
10.8 Conclusion............................ 145
11 A eory of Goal Tracking 147
11.1 TheSimpleView......................... 147
11.2 The Motor Theory of Goal Tracking . . . . . . . . . . . . . . 148
11.3 The Motor Theory and the Teleological Stance . . . . . . . . 151
11.4 TargetvsGoal.......................... 153
11.5 A Dual Process Theory of Goal Tracking . . . . . . . . . . . 155
11.6 PuzzlesSolved? ......................... 157
11.7 Conclusion............................ 158
12 Mind: the Puzzle 161
12.1 AllAboutMaxi ......................... 162
12.2 Infants track false beliefs . . . . . . . . . . . . . . . . . . . . 166
12.3 A Replication Challenge . . . . . . . . . . . . . . . . . . . . 169
12.4 Methodological Defects or Truly Contradictory Responses? . 170
12.5 Models .............................. 173
12.6 The Mindreading Puzzle . . . . . . . . . . . . . . . . . . . . 175
13 ree Levels of Analysis 177
13.1 Tracking Beliefs without Representing Them? . . . . . . . . 177
13.2 Altercentric Interference . . . . . . . . . . . . . . . . . . . . 178
13.3 Mirroring beliefs? . . . . . . . . . . . . . . . . . . . . . . . . 180
13.4 Three Levels of Analysis . . . . . . . . . . . . . . . . . . . . 182
13.5 TaskAnalysis .......................... 184
13.6 Selection and Inhibition . . . . . . . . . . . . . . . . . . . . . 187
13.7 Too Much Mindreading? . . . . . . . . . . . . . . . . . . . . 192
13.8 WhatNow? ........................... 197
vii
14 Mind: a Solution? 199
14.1 Mindreading Is Sometimes Automatic . . . . . . . . . . . . . 200
14.2 Mindreading Is Not Always Automatic . . . . . . . . . . . . 201
14.3 A Dual Process Theory of Mindreading . . . . . . . . . . . . 202
14.4 Speed–Accuracy Trade-Os .................. 204
14.5 What Is a Model of Minds and Actions? . . . . . . . . . . . . 205
14.6 Minimal Models of the Mental . . . . . . . . . . . . . . . . . 207
14.7 Signature Limits in Mindreading . . . . . . . . . . . . . . . . 210
14.8 A Developmental Theory of Mindreading . . . . . . . . . . . 213
14.9 How to Solve the Mindreading Puzzle . . . . . . . . . . . . . 216
14.10 Task Analysis Revisited . . . . . . . . . . . . . . . . . . . . . 218
14.11 Is There Core Knowledge of Minds? . . . . . . . . . . . . . . 219
14.12 Origins of Knowledge of Mind: Rediscovery . . . . . . . . . 219
15 Joint Action 223
15.1 Joint Action vs Parallel but Merely Individual Actions . . . . 224
15.2 SharedIntention......................... 226
15.3 Bratman on Shared Intention . . . . . . . . . . . . . . . . . . 227
15.4 An Inconsistent Triad . . . . . . . . . . . . . . . . . . . . . . 228
15.5 Coordinating Planning . . . . . . . . . . . . . . . . . . . . . 231
15.6 Joint Action in the First Years of Life . . . . . . . . . . . . . 235
15.7 Collective Goals vs Shared Intentions . . . . . . . . . . . . . 239
15.8 Expectations about Collective Goals . . . . . . . . . . . . . . 242
15.9 Conclusion............................ 245
16 Conclusion to Part II 249
16.1 Dual Process Theories . . . . . . . . . . . . . . . . . . . . . . 250
16.2 Pluralism about Models . . . . . . . . . . . . . . . . . . . . . 251
16.3 Goal Tracking Is the Foundation . . . . . . . . . . . . . . . . 252
16.4 When Joint Action Enables Goal Tracking . . . . . . . . . . 253
16.5 Joint Action and the Developmental Emergence of Knowledge 255
Conclusion 259
17 Conclusion 259
17.1 Infants Rely on Minimal Models . . . . . . . . . . . . . . . 260
17.2 As Do Adults, Sometimes . . . . . . . . . . . . . . . . . . 261
17.3 PuzzlesMatter.......................... 262
17.4 Linking Problems Abound . . . . . . . . . . . . . . . . . . . 263
17.5 Core Knowledge Isn’t What You Think It Is . . . . . . . . . . 264
17.6 How to Solve Linking Problems . . . . . . . . . . . . . . . . 265
17.7 Representation: Handle with Care . . . . . . . . . . . . . . . 266
viii
17.8 Inferential and Intentional Isolation . . . . . . . . . . . . . . 267
17.9 Rediscovery Is Joint Action . . . . . . . . . . . . . . . . . . . 268
Glossary 271
ix

The Developing Mind

A Philosophical Introduction

1. Introduction

At the outset, we humans know nothing, or not very much.

Some time later, if things go well, we do know some things.

How does the transition occur? How do humans come to know about objects, actions and minds?

This question belongs to a family of questions about the origins of mind that philosophers have been asking for a while. In a beautiful myth, Plato suggests that the answer is recollection. Before we are born, in another world, we become acquainted with all the truths we will ever know. Then we are involved in an unfortunate traffic accident and fall to Earth, forgetting everything. But as we grow, we are sometimes able to recall parts of what we once knew. So it is by recollection that humans come to know about objects, actions and minds.1

How else could this happen? Since Plato, philosophers and psychologists have offered other stories. Some hold that knowledge is in some sense present at birth, or else that the concepts which make knowledge possible are already present at birth. Others suggest that concepts and knowledge are acquired through sensory experience, through learning to act, through training in language or through social interaction. None of these bold, seductive ideas is supported by much evidence. They are too difficult to test, or perhaps not even precise enough to test. And when we look at particular domains of knowledge in detail—for instance, when we look at how humans come to know about minds—we will discover complexities that seem to be incompatible with any one of the stories. While it would be fun to tour nativism, empiricism and other big ideas about the developmental origins of human knowledge, we are unlikely to make much progress if the last couple of millennia are any guide to the future. Let’s try a different approach.

Start with the details. Take one domain of knowledge—knowledge of objects, say. What has been discovered about infants’ abilities in this domain, and about how knowledge of simple facts in this domain emerges in development? Pursuing this question leads directly to puzzling patterns of evidence. These puzzles in turn point to theoretical challenges requiring, often enough, broadly philosophical solutions. Identify those puzzles and distinguish candidate solutions. In the best case, one of the candidate solution’s predictions will turn out to be largely correct, and we will all have taken a tiny step towards understanding the developmental emergence of knowledge.

This book is a philosophical introduction to how, from earliest infancy, human minds develop and acquire knowledge. Drawing on discoveries in developmental psychology, it aims to introduce readers to findings, concepts and theories needed to explain the developing mind. But it parts company from developmental psychology in that it is written from a philosophical standpoint. As such, we will focus on puzzles that arise in investigations of how knowledge of objects, minds and actions develops. Attempting to solve these puzzles will require us to consider fundamental questions in the study of the mind. These comprise both architectural questions about modularity, core systems and dual process theories, as well as questions about the role of practical and inferential reasoning, mental representation, metacognition, belief, perception, innateness, mindreading and joint action.

1.1 Two breakthroughs

Can developmental philosophical psychology take us further than Plato got? Maybe. Two relatively recent scientific breakthroughs promise to shift thinking away from myths and closer to the minds and actions of actual humans. The first breakthrough concerns social interaction. It is the discovery that preverbal infants enjoy surprisingly rich social abilities. These may well facilitate the subsequent acquisition of linguistic abilities and enable the emergence of knowledge (as variously argued by several people, including, for example, Tomasello et al. 2005; Meltzoff 2007; Csibra and Gergely 2009).

A second breakthrough involves the use of increasingly sensitive—and sometimes controversial—methods to detect expectations without relying on subjects’ abilities to talk or act. These methods have revealed that, from the early months of life on, infants have sophisticated abilities to track physical objects and their causal interactions, actions, mental states and more besides (for example, Spelke 1990; Baillargeon, Scott and He 2010). The mental states underpinning these abilities surely play a role in the emergence of knowledge.

Although each of these breakthroughs has been extensively discussed, they are rarely considered together. There may be an opportunity to make progress by combining the breakthroughs. My guess is that development is like climate change in one respect. Lots of different mechanisms are simultaneously at work, and many interact with each other. To make progress, we need to identify various mechanisms and understand their interactions. This book is an attempt to show, by closely following what has been discovered so far, that understanding the emergence in development of knowledge will eventually require somehow bringing together the abilities that infants manifest in the very first months of life concerning physical objects, minds and actions and their abilities to act jointly with those around them.

Before we get to the details, let me outline a little theoretical background.

1.2 Knowledge

The question we face—How do humans come to know about objects, actions and minds?—is a question about knowledge. Answering this question depends on discovering when humans come to know what. And making these discoveries in turn depends on being able to distinguish really knowing something from merely manifesting some symptoms associated with knowledge.

Imagine an infant who seems to want a toy and, when given the chance, immediately searches for it in exactly the place it was lost. She is acting as if she knew where it was lost, so exhibiting a symptom of knowledge. But does she really know? Maybe not. If the state underlying her searching actions were locked to an arbitrarily limited range of actions, say, then it would not be knowledge. So what is distinctive of really knowing something?

In what follows, I take for granted that knowledge is constitutively linked to practical reasoning and to inference. Let me explain. Knowledge is the kind of thing that can typically influence how you act when you act purposively, and it is the kind of thing that can influence purposive actions in any domain at all. Knowledge is also the kind of thing that you can sometimes arrive at by inference, and which can enable you to make new inferences in any domain at all. A state that is not linked to practical reasoning and inference in these ways is not knowledge.

I also take for granted that knowledge states are inferentially integrated with other attitudes like beliefs, desires and intentions. This does not mean, of course, that people are invariably rational. Instead the idea is this. One striking fact about many humans is that, at times, they achieve a kind of harmony in what they know, believe, intend, desire and do. Sometimes, some of their thoughts and actions come to be approximately rationally related. Now it may be that this is rare, or even highly unusual, in humans. But however infrequent, since it is not an accidental occurrence, it stands in need of explanation. And the explanation, or part of it, involves processes of practical reasoning and inference. In saying that knowledge states are inferentially integrated with other attitudes like beliefs, desires and intentions, part of what I mean is that these they can (albeit perhaps rarely) come to be non-accidentally related in ways that are approximately rational thanks to processes of inference and practical reasoning.

But there is more to being inferentially integrated. When humans are functioning at their best, they characteristically bring thoughts and actions into harmony unless something prevents them. This is the other part of what I mean by saying that knowledge states, beliefs and the rest are inferentially integrated: in the absence of obstacles such as time pressure, distraction, motivations to be irrational, self-deception or exhaustion, approximately rational harmony will characteristically be maintained among currently active knowledge states, intentions and other attitudes.2

These facts about knowledge are almost too simple to mention. But they will turn out to be critical for distinguishing really knowing something from merely manifesting some symptoms associated with knowledge. The hypothesis that someone knows something generates the prediction that, in the absence of obstacles, she can manifest this knowledge in almost any situation.

If it is locked to an arbitrarily limited range of actions, or if it used in response to an arbitrarily limited range of events, then it is not knowledge.

1.3 A crude picture of the mind

Knowledge and the other attitudes contrast with perceptual representations. These are those postulated by scientific theories to explain processes such as edge detection or the computation of relative distances (see Palmer 1999, for an introduction).

Knowledge also contrasts with motor representation, which is less familiar but will be important later. Imagine being in a coffee shop where the servers use little round trays. You are watching the servers as they remove items from trays they are carrying around. As they lift a mug from a tray, the tray remains stable. How do the servers do this? They are not deliberating about the forces involved (or not usually). But nor is this a mindless physiological change. Instead it involves anticipation of the effects of their own actions, as you can discover by removing an item from the tray when a server is not looking—this can easily cause them to drop everything on the tray. Anticipatory control of action is one of things motor representations enable (see Rosenbaum 2010, for an introduction). They are those representations of actual, possible, imagined, or observed actions and their effects which are characteristically involved in preparing, performing and monitoring sequences of small actions such as grasping, transporting and placing a mug.

Unlike knowledge states, perceptual and motor representations are plausibly not inferentially integrated with beliefs, desires, intentions and other attitudes. You can have perceptual experiences of the relative sizes, colours or locations of objects which are incompatible with what you know and believe. Such cases—illusions—are not due to you simply failing to make an inference. Nor are they symptoms of self-deception or a divided mind. They are consequences of the fact that perceptual processes are, to an interesting extent, distinct from the inferential processes in which knowledge states feature. This is why there are illusions, and, more generally, why a single event can result in multiple incompatible representations.3

Let us take as our starting point a crude but quite standard picture of the adult mind. The mind comprises at least three kinds of states and processes:

1 epistemic (that is, knowledge-related); 2 motoric; 3 perceptual.

The three kinds of process are to an interesting extent distinct from each other, and the three kinds of states are not inferentially integrated in the above sense.

When we explore recent discoveries about infants’ abilities, we will see that they do not fit neatly with this crude picture of the mind. They appear to be in states which are not epistemic, not motoric and not perceptual. This will be a key theme in the following chapters: understanding the developmental emergence of knowledge requires identifying states which do not fit neatly into the crude picture of the mind. One of the major unresolved challenges is finding a good way to revise the crude picture, one that can generate novel predictions.

1.4 Core knowledge

The need to identify states which do not fit neatly into the crude picture of the mind has been discussed by Davidson, although in his view the need arises for philosophical reasons rather than as a consequence of any scientific discoveries. He writes:

The difficulty in describing the emergence of mental phenomena is a conceptual problem … In … the evolution of thought in an individual, there is a stage at which there is no thought followed by a subsequent stage at which there is thought. To describe the emergence of thought would be to describe the process which leads from the first to the second of these stages. What we lack is a satisfactory vocabulary for describing the intermediate steps.

(Davidson 2001, 127)

Where Davidson uses the word ‘thought’, I am using ‘knowledge’. This difference is unimportant here (because of inferential integration).

Davidson goes on to say that the problem cannot be solved: ‘If you want to describe what is going on in the head of the child when it has a few words which it utters in appropriate situations, you will fail’ (2001: 127–8). But will we fail? Since describing ‘what is going on in the head of the child’ is the focus of much developmental psychology, perhaps there are some ideas that will help.

One key idea from developmental psychology is that of core knowledge (Spelke et al. 1992; Carey and Spelke 1996; Spelke 2000). As Carey puts it, the central claim is this: ‘there is a third type of conceptual structure, dubbed “core knowledge” … that differs systematically from both sensory/perceptual representation[s] … and … knowledge’ (2009, 10; my emphasis). Core knowledge states feature in core systems, which are ‘largely innate, encapsulated, unchanging, arising from phylogenetically old systems, and built upon the output of innate perceptual analyzers’ (Carey and Spelke 1996, 520). For many domains of knowledge, the best-supported developmental theories postulate the existence of a core system as something more primitive than knowledge. These core systems are thought to provide a basis for the developmental emergence of knowledge.

I emphasized ‘third type’ in the quote just above because core knowledge is supposed to be a state distinct from knowledge (and belief). In terms of the above crude picture, which distinguishes epistemic, perceptual and motor representations,4 we should really say that core knowledge is supposed to be a fourth type of state.

A key question in what follows is whether Spelke, Carey and others are right that explaining development requires postulating a third or fourth type of mental state, something which is not epistemic, perceptual, or motoric. This question is linked to a long-standing conflict between two stories about development.

1.5 Two stories

Any story has to accommodate the breakthrough discovery that infants, even in their first months, have sophisticated abilities to track objects, causal interactions, numerosity, actions, mental states and more besides.

Perhaps the theoretically simplest (in a good sense) way to do this is to invoke knowledge. According to one story, infants’ earliest abilities to engage with objects, actions and mental states are based on knowledge. From as early as they can manifest these abilities, they know some general principles in whatever sense the adults do. They use this initial knowledge of general principles together with perceptual information to make inferences about particular things—objects, actions, minds and the rest. Infants differ from adults only in that they do not know very much, whereas, when things go well, adults know more. Development is essentially just a matter of acquiring more knowledge.

A strong case for this first story can be made, as we will see for one domain of knowledge in Chapter 3. In essence, hypothesizing that infants know certain things enables us to characterize their abilities in a way that is both theoretically simple and mostly accurate. The problem is that such hypotheses also systematically generate incorrect predictions.

The other story is about core knowledge. On this story, infants’ earliest abilities to engage with things are based on core knowledge of general principles (rather than on knowledge proper). This core knowledge somehow provides a basis for the acquisition of knowledge. And when infants do succeed in acquiring knowledge, this does not change their core knowledge. Instead their core knowledge is constant through life, even if it conflicts with things they later come to know. So in many domains there are two types of state that can influence cognition and behaviour, core knowledge and knowledge proper.

The second story aims to accommodate the breakthrough discovery about infants’ early abilities while avoiding the incorrect predictions generated by conjectures about knowledge. Accordingly, the case for this second story is based on both evidence used to support the first story and evidence against it.

Although the second story does have the advantage of not generating incorrect predictions, it does have another weakness. As core knowledge is usually characterized, this story fails to generate relevant predictions (see Chapter 5).

At this point we reach an impasse. The challenge is to characterize what is going on in an infant’s mind, and what is driving her actions, before she has any relevant knowledge states or beliefs. Neither of the standard stories seems adequate. One is simple and predictively strong but generates incorrect predictions; the other is theoretically complex but predictively weak. What to do?

1.6 Development is rediscovery

Here is a preview of my story. It starts from the idea that core knowledge is not one thing: it lacks both unity and uniformity. The phenomena associated with core knowledge are not exclusively either perceptual or motoric, although some are broadly perceptual and some are broadly motoric. And they comprise different kinds of representations and cognitive structures in different domains.

If this is correct, current attempts to provide general theories of core knowledge are misguided because they assume unity (core knowledge is one thing) and uniformity (core knowledge is the same thing in different domains). Fortunately such theories are also mostly unnecessary. We can make good progress in understanding ‘core knowledge’ (or whatever else you would like to call infants’ earliest cognition of objects, actions, minds and the rest). This is because infants’ abilities and their limits can sometimes be explained by representations, structures and processes which have been identified and studied independently of developmental theories. As we will see, in some domains, there is evidence to support detailed conjectures from which novel predictions flow (see Chapters 6, 7, 10 and 14).

Thinking about core knowledge in this way motivates construing development as a process of rediscovery. Infants’ sophisticated abilities to track objects, actions and mental states reveal that, even in the first months of life, their cognition and action are influenced by an impressive range of truths and useful falsehoods about the natures of these things. (This is the breakthrough discovery that led to theories about core knowledge, of course.) But these truths and falsehoods, if represented at all, are represented in ways that are cut off from knowledge. The representations are not inferentially integrated with knowledge and cannot lead to the acquisition of knowledge by inference. Further, their causal interactions with knowledge may involve intermediaries which either lack intentional features entirely or else have intentional features that are only very distantly related to those of the cognitive structures they link. That is, they may be intentionally isolated from knowledge.

Imagine a single organization with different parts. One part has information about criminals which enables it to predict and detect break-ins. This is the security department. Another part of the organization lacks this information. But instead of communicating internally, this part goes out and learns everything from scratch; and what it learns is never fed into the security department. This is the research department. From the point of view of this department, it is almost as if the organization starts with no knowledge of criminality at all. But only almost—for the learning done in the research department depends on the operations of the security department to keep it safe, and to flag up criminal events. This is rediscovery: one part goes out and learns afresh about things related to those already relied upon by another part of the organization; meanwhile the other part’s ongoing contribution to the organization’s flourishing continues unaffected by any new learning.

You might think this analogy shows that the idea of rediscovery is flawed. After all, this kind of organizational behaviour is routinely condemned as failure. But primate minds are special kinds of organization. They enable rapid learning in ways which are open to radical revision and can therefore yield radically mistaken conclusions. This learning should not be contaminated by assumptions which happen to enable the animal to function from the first months of life onwards, and which are effective even when it has little experience of the world and limited social interactions. And nor should the mistakes it makes in learning compromise its ability to engage effectively with the physical and social aspects of the world. Both requirements can be met thanks to the inferential and intentional isolation of the security department from the research department—or of the representations involved in early-developing abilities from knowledge.

If inferential and intentional isolation mean that development is rediscovery, we face a challenge. How could anything like core knowledge ever facilitate the developmental emergence of knowledge?

The phenomena associated with core knowledge and the infant’s knowledge and beliefs can be connected indirectly, via her behaviour, experience and attention. To illustrate with a comparatively simple example, consider the development of face perception. From birth, there is a mechanism present in humans (and chickens) that uses crude heuristics to identify faces and generates orienting reflexes. There is also a later-developing mechanism which uses more sophisticated, learned principles geared to features of conspecifics and enables smooth tracking of moving faces. Crucially the two mechanisms do not share any representational resources; the second mechanism does not rely on signals carrying information about whether the first mechanism has detected a face. Instead they are linked only indirectly, via behaviour: the first orients the baby’s eyes to faces, thereby indirectly providing inputs necessary for learning in the later-developing mechanism (Johnson and Morton 1991; de Haan 2002). This is one illustration of how, despite their intentional isolation, an early-developing ability can facilitate the emergence of a later-developing one indirectly, by influencing behaviour, experience and attention.

The case of face detection also shows, by the way, that arrangements involving multiple mechanisms are not necessarily wasteful. In this case the arrangement makes sense, given the different needs humans have at different times of life. The first mechanism relies on fixed heuristics that work well in the particular position that new-born infants, who cannot support their heads, mostly find themselves in. This enables infants to reliably identify, and orient to, faces from birth. Orienting to faces provides useful input for the second, later-developing and more flexible mechanism. The first does not tell the second which things are faces or how to detect them: they are intentionally isolated. And because it is not constrained by the fixed heuristics of the first mechanism, the second mechanism eventually enables older infants to detect faces in a wider range of situations. Rediscovery costs time and cognitive effort, of course. But it provides benefits in accuracy and flexibility.

One crucial feature is missing from this illustration. If development is rediscovery, and if rediscovery is something achieved by acting together with those who care for us, then infants’ social skills are indispensable drivers of knowledge. That is why Part I of this book focuses on infants’ abilities concerning physical objects and Part II on their social skills. I will not be attempting to construct an account of how knowledge emerges in development; but I will introduce the two breakthroughs that, eventually, must somehow be combined in explaining the developmental origins of knowledge.

Notes

Part I

Physical objects

2. Principles of object perception

How do humans first come to know simple facts about particular physical objects? To illustrate, you probably know some facts about the approximate location and shape of the book you are reading or the device you are using to read it. But no one is born knowing any facts about particular physical objects. So there was a time when you knew no such facts about any particular physical objects at all, and then, some time later, you came to know some such facts. What makes this amazing transition possible?

This chapter introduces the experiments and discoveries that will guide our thinking in working out how humans first come to know facts about particular physical objects. Another aim of this chapter and Chapter 3 is to introduce some methods most commonly used in research on human infants. By the end of these two chapters you should be familiar with experiments involving habituation, violation-of-expectation, anticipatory looking and search behaviours. (Methods involving neuroscience will not concern us until later chapters.)

2.1 Knowledge of objects involves three abilities

What does knowledge of physical objects involve? To answer this question, it is helpful to think about features common to all physical objects.

One feature is boundedness: there is a fact of the matter about where one physical object ends and another begins. Since physical objects have boundaries, knowledge of physical objects plausibly involves an ability to identify where one object ends and another begins—that is, to segment them. As Figure 2.1 illustrates, the ways objects are ordinarily arranged in space prevents us from doing this in any very simple way. Sometimes one object occludes parts of another from your point of view, and sometimes one object contains another. So distinct objects do not always occupy clearly separate regions of space. To count the number of geese in a scene, or even just to distinguish which of several geese-parts belong to the same goose, involves being able to segment objects.

Figure 2.1

Figure 2.1 The ways objects are ordinarily arranged in space means that there is no simple way to segment them

Knowing where to find her and wondering whether this girl is her both involve being able to represent objects as persisting.

Physical objects can causally interact with each other. Knowing any facts about physical objects is likely to involve an ability to track causal interactions. Minimally, this involves being able to distinguish interactions from mere contingencies. Seeing how well the deformed shape of the barrier matches the dent in a nearby car, you might wonder whether this is mere coincidence or a consequence of the car hitting the barrier. Seeing how one moving ball stops just as another begins to move, you might guess that the second ball’s movement is caused by that of the first. In doing these things you are relying on an ability to track causal interactions among objects.

So far we have seen that reflection on features common to all physical objects suggests that knowing even simple facts about physical objects plausibly involves three abilities: (1) to segment objects; (2) to represent them as persisting; and (3) to track some of their causal interactions. Our main question is, How do humans first come to know any facts about particular physical objects? We can approach this question by asking a simpler one: When and how do humans first manifest the three abilities associated with knowledge of objects?

2.2 Segmentation

How do humans first segment physical objects? Textures and patterns provide one possible clue to where boundaries between objects lie. As mentioned earlier, because one object often touches, occludes or contains parts of another object, objects do not reliably occupy clearly separate regions of space. But a sharp change in texture is an indicator—not decisive but perhaps reliable enough—of a boundary between two objects. Consider the top left panel in Figure 2.2. The abrupt change in texture indicates that there are two similarly sized objects,

Figure 2.2 Stimuli from an experiment showing that 4½-month-olds can use information about patterns or textures to segment objects

Figure 2.2

Source: Needham (1998), Figure 3.

one on its side and the other standing up. That adult humans can use texture to segment objects is hardly controversial, of course. But when in their development can humans first do this?

To answer this question, Needham devised a simple test (Needham 1998, experiment 2). She first showed her subjects, who were 4½-month-old infants, a display like that depicted in the top left panel of Figure 2.2. A hand then appeared, grasped part of a visible object and pulled it. For some of the infants, the hand’s pulling resulted in the moving apart of two objects, as depicted in the top row of pictures in Figure 2.2. For the other infants, the hand’s pulling resulted in everything moving together, as depicted in the bottom row of pictures. Would one group of infants look at the display for significantly longer than the other group? Needham reasoned that different hypotheses about segmentation predict different answers to this question. Consider the hypothesis that infants do use textural information to segment the objects. If this is true, infants would represent the scene as containing two objects and so would expect these objects to move independently when a hand pulls one of them. Given that infants generally look longer at things which violate their expectations, those infants who see the two objects moving together should look at the display for longer than those who see the objects moving apart. By contrast, the hypothesis that infants do not use textural information to segment objects does not generate any such prediction about relative looking times. Accordingly, finding that those infants who see everything moving together look longer than those who see objects moving apart would be evidence that infants can use textural clues to segment objects. And this is just what Needham found, as you can see in Figure 2.3.

Needham’s experiment is relevant not only as evidence that infants can use texture to segment objects from their first months of life. It also illustrates a widely used method that will come up frequently. The method is called violation-of-expectation. It relies on the premise that when something violates an infant’s expectation, she will typically look at it for longer than she otherwise would have. In designing a violation-of-expectation experiment, you first

Figure 2.3 Infants can use texture to segment objects from around 4½ months of age

Source: Needham (1998), Figure 4.

identify a hypothesis which generates predictions that certain individuals have particular expectations. (In Needham’s experiment, the hypothesis was that 4½-month-olds use texture to segment objects; this generated predictions about when they would expect things to move apart.) Importantly, no other relevant hypothesis should generate these predictions. You then contrive two scenarios which are as similar as possible except that the putative expectation is violated in one but not in the other scenario. If you find that subjects look significantly longer at the scenario in which the putative expectation is violated, you have discovered evidence for the hypothesis. With careful use of the violation-of-expectation method, infants’ looking times enable us to discover many surprising things about what they expect.

Talk of expectations being violated may leave you wondering, What is an expectation? This is an important question, and a surprisingly deep one. We will return to it in Section 8.1.

The main question for this section is how humans first segment physical objects. We have seen that infants in their first months of life can use texture to segment objects; further research also shows that they can also use shape as a clue (Needham 1999). But do humans’ abilities to segment objects initially rely entirely on superficial aspects of objects like their textures and shapes?

To show that they do not, Kellman and Spelke (1983) asked what happens when infants are shown a stick partially occluded by a box, as illustrated in Figure 2.4. The stick is moving while the box remains still. How do infants represent this scenario? Does it look to them as if there are three objects, two short sticks either side of a box? Or do they represent the scenario as most adults would—that is, as involving one long object partially occluded by the box? To answer this question, Kellman and Spelke (1983) showed a group of 4-month-old infants the stick moving behind the box repeatedly, until it no longer held their interest. (That is, the infants were habituated to the display.) After this, some of the infants were shown a new scenario involving a single stick moving; and the other infants were shown a different new

Figure 2.4 A stick moving behind a static box

Source: Kellman and Spelke (1983), Figure 3, part.

scenario involving two unconnected sticks moving simultaneously (see Figure 2.5). In both new scenarios, the visible parts of the stick from the first scenario were unchanged. How different are each of the new scenarios from the original scenario? That depends on how you represent the original scenario. If you represent the original scenario as containing one stick partially occluded by a box, then the new scenario with two sticks is more novel than the new scenario with just one stick—not only has the box gone, but now there are two sticks whereas before there was just one. (By saying that one of the new scenarios is ‘more novel’, I mean that it differs more from the original scenario than the other new scenario differs from the original scenario.) But if you represent the original scenario as containing two sticks either side of a box, then the converse is true: the new scenario containing just one stick is more novel. So if we had some way of finding out which of the two new scenarios infants find more novel, then we could work out whether infants represent the original scenario as involving one stick or two. But finding out which of the new scenarios infants find more novel is surprisingly simple (in theory at least—little involving infants is simple in practice). In general, after infants have been habituated to one scenario and shown new scenarios, the more novel new scenario will produce greater dishabituation. That is, the more novel scenario will be more effective in regaining infants’ interest.101 Accordingly, Kellman and Spelke measured which of the two stick scenarios depicted in Figure 2.5 produced greater dishabituation. You can see their results in Figure 2.6. The left part, labelled ‘habituation’, describes how infants spent less time looking at the original scenario after seeing it repeatedly. The first data points on the right part, labelled ‘test’, reveal a sharp increase in looking time for those infants shown the two-stick scenario (labelled ‘broken’) contrasting with no discernible increase in looking time among those infants who saw the one-stick scenario (labelled ‘complete’). That is, the two-stick scenario produced greater dishabituation. This is evidence that this scenario is more novel to the infants, which indicates that infants represent the original scenario as containing a single stick partially occluded by a box.

Figure 2.5 Two scenarios, both possible consequences of removing the box from Figure 2.4

Source: Kellman and Spelke (1983), Figure 3, part.

Figure 2.6 Evidence that 4-month-olds segment objects using cues other than feature and shape

Source: Kellman and Spelke (1983), Figure 4, part.

What is exciting about a stick moving behind a box? The fact that infants represent this scenario as involving one stick rather than two is a hint that they are not segmenting objects just on the basis of superficial properties like shape or texture. This idea is supported by evidence from a further experiment. Kellman and Spelke modified their stick-moving-behind-box apparatus so that the visible portions of the object moving behind the box were markedly different in shape and texture, as depicted in Figure 2.7: and they made a corresponding change to what infants saw in the test phase of the experiment. Although the two parts of the moving object are so different, 4-month-olds’ patterns of dishabituation again showed that they represent the scene as involving a single, connected object behind a box (Kellman and

Figure 2.7 An object with parts that differ in shape and texture moving behind a box

Source: Kellman and Spelke (1983), Figure 13, part.

Spelke 1983, Experiment 6). Taken together, these experiments provide good evidence that infants’ abilities to segment objects are not based entirely on recognizing shape and texture (see also Spelke 1990).

2.3 Principles of object perception

If infants do not rely only on shape and texture, how do they segment the objects in the displays we have just been considering? Spelke (1990) suggests that infants rely on a set of principles to segment objects.

One of her principles is no action at a distance. This principle says that distinct ‘objects are interpreted as moving independently of one another’.102 Because the two moving parts of the apparatus represented in Figure 2.7 do not move independently, the principle indicates that they are not distinct objects. So the hypothesis that this principle describes in part how infants segment objects correctly predicts that they will treat the moving occluded stick as a single object.

The fact that the principle of no action at a distance describes how infants segment objects in one case (the stick-moving-behind-a-box case) is interesting but not by itself terribly informative. What is more interesting is that this principle describes how infants segment objects in a wide variety of cases (see Spelke 1990, for a more comprehensive review). To get a feel for this variety, consider another study involving different stimuli and a different measure—instead of looking times, this one uses infants’ reaching actions.

Let me explain the stimuli first. The study involves the four scenarios depicted in Figure 2.8. In the scenario depicted in the bottom right of this figure, two spatially separated regions move rigidly in the same direction (consistently with their being parts of a single, partially obscured object). The no-action-at-a-distance principle implies that there is just one object in this scenario. Now consider the scenario depicted in the top left of Figure 2.8. In this scen-

Figure 2.8 Four scenarios. Things move in opposite directions in the two scenarios depicted left, whereas in the other scenarios everything moves together.

Source: Spelke, Hofsten, and Kestenbaum (1989), Figure 3.

ario, two spatially contiguous things are moving in opposite directions. Intuitively, these movements imply that there are two objects in this scenario. This intuition is codified in a further principle Spelke calls rigidity. This principle is the converse of the no-action-at-a-distance principle; it states that ‘objects are interpreted as moving rigidly’ (Spelke 1990, 50). As there is no way of interpreting the scenario in such a way that one object moves rigidly, the principle of rigidity implies there are two objects. Spelke, Hofsten and Kestenbaum (1989) set out to show that, in each scenario depicted in Figure 2.8, infants’ reaching behaviours indicate that they segment objects in accordance with the principles of no action at a distance and rigidity.

How can we infer any such thing from reaching? Spelke, Hofsten and Kestenbaum (1989) set things up so that the smaller of the two objects was always closer to the infant. The experimenters were then able to rely on the background assumption (which is not obvious but carefully justified, see Spelke, Hofsten and Kestenbaum 1989, 186) that infants typically reach more often for the smaller, nearer object when they represent the scenario as involving two objects than when they represent it as involving just one object. So by comparing how often 5-month-olds reach for the smaller object, Spelke, Hofsten and Kestenbaum (1989) could determine, for each scenario, whether infants treat it as involving two objects or just one. As predicted, they found that, overall, 5-month-olds reach for the smaller, nearer object more often when they moved in opposite directions than when they moved together. Given the background assumption, this is evidence that infants segmented the objects in accordance with the principles of rigidity and no action at a distance.

No action at a distance and rigidity are not the only principles we need to explain how infants segment objects. Other principles which seem to be involved in segmenting objects are cohesion and boundedness. According to the principle of cohesion, ‘two surface points lie on the same object only if the points are linked by a path of connected surface points’ (Spelke 1990, 49). The principle of boundedness says, conversely, that ‘two surface points lie on distinct objects only if no path of connected surface points links them’ (Spelke 1990, 49). To illustrate, consider the two situations represented in the left and right parts of Figure 2.9. The principle of boundedness implies that there is one object in the left-depicted situation, and the principle of cohesion implies that there are two objects in the right-depicted situation. Kestenbaum, Termine and Spelke (1987) suggest that these two principles are also needed to describe how infants segment objects.

Figure 2.9 According to the principles of cohesion and boundedness, the situation depicted on the left involves one object whereas the situation depicted on the right involves two

Source: Kestenbaum, Termine and Spelke (1987), Figure 3, part.

2.4 Conclusion

Our ultimate aim is to understand the transition humans make from knowing nothing at all about particular physical objects to having such knowledge. Given how difficult this is, we have taken an indirect route by first considering three abilities that are probably involved in having knowledge: abilities to segment physical objects, to represent them as persisting and to track their causal interactions (see Section 2.1). The latter two abilities are a topic for Chapter 3; this chapter was about how humans first segment physical objects.

Before looking at the evidence, it might have been tempting to guess that very young infants cannot segment objects at all, or else that they can only segment objects by using surface properties like shape and colour. The truth turns out to be much more interesting. From around 4 months of age, humans can indeed use shape and texture in detecting where one object ends and another begins. But by this age they also segment objects in accordance with a set of principles which go beyond the surfaces to describe the characteristic ways objects move. As we have seen, one consequence is that infants, like adults, can reliably form expectations about occluded parts of objects. In fact, it seems that, from as early as they have been tested, infants segment objects in much the way that adults do.

This is an exciting result, but it does not yet enable us to answer the question we started with. The question is, How do humans first segment physical objects? We can say that they do it in accordance with the four principles that Spelke labels Principles of Object Perception—no action at a distance, rigidity, boundedness and cohesion. But to say this is merely to describe patterns in their performance. Answering our ultimate question—How do humans first come to know simple facts about particular physical objects?—will require discovering what explains these patterns.

Notes

3 The Simple View

We have just seen in Chapter 2 evidence for Spelke’s bold conjecture that humans, from the first months of life, segment objects in accordance with the Principles of Object Perception. This is a bold conjecture: it cites a small number of principles which generate many testable predictions about how infants segment objects in different situations. The success of these predictions (so far, at least) motivates us to ask two questions. First, what is the link between these principles and the mind of an individual infant? Second, are whatever abilities 4-month-olds have to represent objects as persisting or to track their causal interactions also described by the same principles? A positive answer to this second question will hint that infants’ abilities concerning physical objects may all be consequences of a single system. Answering the first question will be the key to characterizing this system. In answering these questions we are aiming to identify the basis for infants’ earliest knowledge of particular physical objects, so approaching our ultimate goal of understanding the transition from knowing nothing about particular physical objects to knowing something.

3.1 The Simple View

What is the link between the Principles of Object Perception which describe 4-month-olds’ abilities to segment physical objects and the mind of an individual infant? Start with an answer I shall call the Simple View: The Principles of Object Perception are things that we believe or know, and we acquire beliefs or knowledge about particular physical objects from these principles by a process of inference. How does this process of inference work? Perceptual processes provide information about the locations of surfaces in space. This information, together with the Principles of Object Perception, provides constraints on what physical objects there are. We form beliefs about particular objects, their locations, movements and interactions in such a way that if our beliefs were true, all of the constraints would be met.

On the Simple View, coming to know facts about objects is a bit like solving a sudoku or crossword puzzle. You start out with some general constraints and a little information. You then try to fill in the gaps in a way that satisfies the constraints.

In what senses is the Simple View simple? It invokes only states and processes that are already familiar in characterizing adults’ abilities, namely, knowledge, belief and inference. So it involves no conceptual novelty. And it specifies a simple relation between infants’ and adults’ abilities, namely, identity.

The Simple View is worth considering because it is so, well, simple. And although Spelke herself would probably no longer endorse it, she does appear to have accepted the Simple View at one point in her thinking: ‘objects are conceived: Humans come to know about an object’s … boundaries … in ways like those by which we come to know about its material composition or its market value’ (Spelke 1988, 198). As we will see in Section 3.3, several other researchers also appear to endorse the Simple View.103

The Simple View is inspired by two famous cognitive scientists, Marr and Chomsky. Marr, building on Helmholtz’ and others’ ideas, showed that many visual processes can be described as inferences (see Marr 1982). And Noam Chomsky pioneered the idea that humans’ knowledge of language depends on their knowing of a small number of principles (see Chomsky 1965). Similarly, the Simple View implies that human infants (and adults) come to know facts about particular physical objects by virtue of making inferences from a small number of principles which they know or believe.

The Simple View matters for understanding how humans first come to know facts about particular physical objects. If the Simple View is right, segmenting objects involves coming to know facts about them. So, 3- and 4-month-olds already know some facts about physical objects (specifically, about where one ends and another begins). And humans do not have an ability to segment objects that is developmentally and conceptually more primitive than their ability to know facts about physical objects. Rather, to segment objects is to know some facts about them. Or so the Simple View implies.

So far the case for the Simple View is based entirely on findings about how humans first segment objects. Having an ability to segment objects (to determine where one ends and another begins) was only one of three abilities associated with knowledge of objects that we identified. It is now time to turn to the second of these, the ability to represent objects as persisting. Thinking about persistence will strengthen the case for the Simple View. It will also bring us one step closer to understanding how humans first come to know facts about objects.

3.2 Persistence

Our overall aim is to discover how humans first come to know simple facts about particular physical objects. Because this is so difficult, we are approaching it by considering three abilities associated with knowledge of particular physical objects (as explained in Section 2.1).

We covered the first, the ability to segment objects, in Section 2.2. The next is the ability to represent things as persisting.

When you are looking after a young child, your representations of the child’s location, size and other simple properties do not depend on your continually perceiving the child. Even if the child momentarily disappears from view, you may well continue to represent her approximate location so that, for instance, you would be surprised to find her sneaking up behind you. Without any ability to represent things as persisting, your world would be limited to the objects you are currently perceiving. Let the child hide behind the logs and, from your point of view, she would cease to exist. Turn around and the logs would cease to exist too. When and how are humans first able to represent objects as persisting?

To answer this question, Spelke et al. (1995) took two groups of 4-month-old infants. The infants in one group were habituated to a display involving continuous movement which adults would naturally describe as a single object moving behind two barriers (see the top left panel of Figure 3.1). Infants in the other group were habituated to a display involving discontinuous movement which adults would typically describe as one object disappearing behind a barrier followed by another object appearing from behind a second barrier (as depicted in the top right panel of Figure 3.1).

Why is this significant? Imagine for a moment that you lacked any ability to represent objects as persisting. Then the two displays would differ only in that there is a short interval without anything moving in the second display. But since you are not representing anything as persisting, from your point of view, the two displays would not differ with respect to the number of objects involved. By contrast, if we think in terms of persisting objects, then the two displays differ in a more substantial way: whereas the first involves a single moving object, the second must involve two moving objects. The hypothesis that 4-month-olds can represent objects as persisting (even while unperceived) therefore implies that infants should represent the continuous movement as involving one object and the discontinuous movement as involving two.

How can we tell whether this is so? Take those infants who were habituated to the continuous movement. After habituation, each was shown either a single moving object (as depicted bottom left in Figure 3.1) or two moving objects (as depicted bottom right in Figure 3.1). Because dishabituation is stronger for more novel displays, the hypothesis that the infants can represent objects as persisting predicts that these infants should show greater dishabituation to the two moving objects. After all, they were habituated to one object moving and are now seeing two. And, by similar reasoning, the hypothesis also predicts that those infants habituated to the discontinuous movements should show the opposite pattern of dishabituation: that is, after their habituation they should look longer if shown one object moving than if shown two objects moving. And this is just what Spelke et al. (1995) found. (Aguiar and Baillargeon (2002) extended this study and obtained convergent findings.) Since the hypothesis that infants cannot represent objects as persisting and live in a world of mere features does not make these predictions, this is evidence that infants, from 4 months of age or earlier, can represent objects as persisting even while they are momentarily out of view behind a barrier.

This conclusion is also supported by experiments involving anticipatory looking rather than habituation. Rosander and Hofsten (2004) showed infants an object moving along a path. At some point it disappeared completely behind a barrier and, predictably, reappeared again as it reached the far side of the barrier. In this situation adults will show anticipatory looking. That is, they will not usually be tracking the object with their eyes while it is fully hidden, but they will position their eyes to just where the object will be at almost exactly the time it should emerge from the far side of the barrier. Rosander and Hofsten (2004) found that, from around 4 months of age, infants show similar anticipatory looking (see also Bertenthal, Gredebäck and Boyer 2013). Their timing is not quite as good as adults’ (it improves with age), but the pattern of anticipatory looking is compelling evidence that even 4-month-olds are in some sense representing objects that are fully hidden from their view as persisting.

How could infants be doing this? Spelke et al. (1995) invoke the Principles of Object Perception, adding a further principle, continuity. According to the principle of continuity, ‘an object traces exactly one connected path over space and time’ and objects’ paths cannot cross (Spelke et al. 1995, 113). Why is this principle relevant to describing how infants represent unperceived objects in the experiments just discussed? Consider again the discontinuous motion depicted top right in Figure 3.1. To interpret this scenario as involving one object would require interpreting that object’s movement as not tracing a connected path (after disappearing behind one barrier, the object would be nowhere until it reappears from behind the other barrier), and therefore as violating the principle of continuity. So supposing that humans, including infants, represent persisting objects in accordance with this principle implies, correctly, that they will interpret the discontinuous motion scenario as involving more than one object.

Figure 3.1 Can 4-month-olds represent objects as persisting? Spelke et al. (1995) habituated one group to the continuous event (depicted top left), and another to the discontinuous event (depicted top right). Those habituated to the continuous event showed greater dishabituation to the two-object display (depicted bottom right) than to the one-object display (depicted bottom left); those habituated to the discontinuous event showed the converse pattern of dishabituation.

Source: Spelke et al. (1995), Figure 2, part.

There is a second way the principle of continuity could in principle be violated. Rather than having one object trace a discontinuous path through space and time, we could have two objects whose paths through space and time cross. This would result in two objects briefly occupying the same space at the same time. Call this a ‘solidity violation’ of the principle of continuity. If the principle of continuity describes how 4-month-olds represent objects as persisting, then they should also react to solidity violations of the principle of continuity. But do they?

Imagine that you are about to walk into a castle by crossing its moat. Just before you get to the drawbridge over the moat, it is drawn up through 90 degrees so that the formerly horizontal bridge is now vertical; it then continues rotating in this direction so that the bridge ends up horizontal again but no longer spans the moat. What you have just observed is a large-scale version of the scenario to which Baillargeon (1987) habituated 4-month-old infants. This is depicted in the top part of Figure 3.2 After infants have been habituated to this scenario, a tiny change is made. Everything is just as before except that before the drawbridge moves, an object is placed behind it. Infants can see the object being placed, but once it is in position, it is hidden from them by the drawbridge. One group of infants then sees the drawbridge rotate through 180 degrees exactly as before, as depicted in the middle of Figure 3.2. This is impossible given that there is a solid object behind it. But if infants do not represent objects as persisting, they should be entirely unaware of this impossibility. The other group of infants were shown the scenario depicted bottom in Figure 3.2. In this scenario the drawbridge stops earlier, at 112 degrees, as if it were impeded by the box; call this the ‘possible scenario’. Baillargeon (1987) asked, Which scenario will cause greater dishabituation? If infants do not represent objects as persisting, then they should be unaware of the impossibility of the drawbridge moving through an object. In that case, the possible scenario should be more novel to them than the impossible scenario because, from their point of view, it differs more from the habituation scenario—something that formerly rotated through 180 degrees is now stopping earlier. By contrast, if infants can represent objects as persisting, then the converse is true: they should find the impossible scenario more novel because this involves something like a magic trick whereas in the possible scenario the movements of the drawbridge are just as expected. Across three experiments, Baillargeon (1987) found that 4-month-old infants showed greater dishabituation to the impossible scenario. This is further evidence that they can represent objects as persisting even while hidden from view.

Baillargeon’s experiments have been the topic of much discussion. Some have replicated her findings (Durand and Lécuyer 2002), and even found related effects with dogs rather than infants (Pattison et al. 2010). But others have attacked Baillargeon’s drawbridge study on methodological grounds. Sirois and Jackson (2012) ran a version of her experiments without finding evidence that infants represent objects as persisting even at 10 months of age. They argue that the apparent novelty of the impossible scenario to infants may be an artefact of the statistics Baillargeon used. As controversies like this abound, it is worth taking a moment to ask how we should respond to it. Can we hold on to Baillargeon’s conclusions or should they be rejected?

Sirois and Jackson’s (2012) methods are rigorous and their criticisms are initially persuasive, so it may be tempting to think that the experiments they target must be dismissed. But things are rarely so straightforward. Sirois and Jackson used computer-generated stimuli

Figure 3.2 Side-on views of the drawbridge scenarios in an experiment showing that 4-month-olds represent objects as persisting

Figure 3.2

Source: Baillargeon (1987), Figure 1, part.

whereas Baillargeon had a physical set-up, they studied 10-month-olds rather than 4-month-olds, and they used a different method (‘children were … not habituated by the time testing began’). So what can we conclude from the fact that Sirois and Jackson (2012) did not find evidence for an ability to represent objects as persisting? This certainly justifies caution in relying on any single experiment. Taken alone, Baillargeon’s (1987) studies are inspiring but not fully convincing. However, many further experiments involving different groups of researchers, different scenarios and different methods provide converging evidence for the same conclusion: even 4-month-olds can represent objects as persisting (for reviews, see Spelke and Hespos 2001 or Baillargeon 2002). The initial, ground-breaking studies are probably methodologically imperfect, but the balance of evidence from subsequent experiments suggests that the discovery they illuminate is probably real.104

Baillargeon’s (1987) drawbridge experiments strengthen the case for saying that infants represent objects as persisting in accordance with the principle of continuity. They suggest that 4-month-olds are sensitive to solidity violations of this principle; that is, to violations which involve two objects briefly occupying the same space at the same time.

The question for this section was: When and how are humans first able to represent objects as persisting? The ability to represent objects as persisting is strikingly similar to the ability to segment objects. Both abilities appear early in development, being clearly present by 4 months of age or earlier. And both abilities can be characterized, at least partially,105 by a small number of principles, namely the Principles of Object Perception. So although abilities to segment objects and to represent them as persisting are conceptually distinct, it may turn out that, in humans at least, there is really just one ability.

3.3 Extending the Simple View to persistence

We are far from fully understanding how humans are first able to represent objects as persisting, of course. But the fact that the ability appears so early in development entails that it does not demand language, nor much conceptual sophistication. This view is supported by the fact that the ability to represent objects as persisting is found in a wide variety of nonhuman animal including monkeys (Santos, Seelig and Hauser 2006), lemurs (Deppe, Wright and Szelistowski 2009), dogs (Kundey et al. 2010), wolves (Fiset and Plourde 2013), cats (Triana and Pasnak 1981), crows (Hoffmann, Rüttler and Nieder 2011), chicks (Chiandetti and Vallortigara 2011) and dolphins (Jaakkola et al. 2010).106 It is possible that humans’ abilities to represent objects as persisting are unrelated to some or all of these other animals’ abilities, of course. Nevertheless, the fact that chicks can represent objects as persisting does show that doing this is not necessary something that requires much cognitive effort or conceptual sophistication.

Can we say more about how humans first represent objects as persisting? The position we are currently considering is the Simple View (which was introduced in Section 3.1. According to the Simple View, there is a set of principles about physical objects and their movements, the Principles of Object Perception; these principles are things that we know or believe; and we generate expectations from these principles by a process of inference. Earlier, in Section 3.1, we saw that the Simple View provides a candidate explanation of humans’, including infants’, abilities to segment objects. But it now seems plausible that a single set of principles will characterize both how humans segment objects and how they represent them as persisting (as we saw in Section 3.2). If so, the Simple View also provides a candidate explanation of how humans represent objects as persisting. They believe or know certain principles and they form expectations about the number and location of objects by making inferences from these principles.

The Simple View is quite widely endorsed. Commenting on findings about infants’ abilities to represent objects as persisting, Aguiar and Baillargeon write: ‘To make sense of such results, we … must assume that infants, like older learners, formulate … hypotheses about physical events and revise and elaborate these hypotheses in light of additional input’ (2002, 329).

If you are inclined to doubt the Simple View—perhaps you doubt that 4-month-old infants can really have beliefs or make inferences—you might suppose that this talk about formulating and revising hypotheses should not be taken literally. You might look for an interpretation on which Aguiar and Baillargeon (2002) do not really mean that infants formulate, revise and elaborate hypotheses. Could they be invoking some kind of tacit or implicit state, and some kind of inference-like process which falls short of actually being inference? We will indeed eventually consider views along these lines in Chapter 5. But for their part, Aguiar and Baillargeon (2002) explicitly specify that infants formulate and revise hypotheses ‘like older learners’. This suggests that their position is probably better captured by the full-fat Simple View than by invoking some semi-skimmed, not-knowledge-but-a-bit-like-it state.

It is also worth noting that merely stepping back from knowledge by invoking an unspecified notion of tacit or implicit knowledge would not amount to providing an alternative to the Simple View. A genuine alternative to the Simple View needs to identify which states and processes are linking an individual’s mind to the Principles of Object Perception. As even articulating (let alone defending) a genuine alternative to the Simple View turns out to be surprisingly difficult, we should hold on to it as our working hypothesis until we encounter evidence against it.

3.4 Causal interactions

Anyone who knows even the simplest fact about a particular physical object—knows its location, say—can probably segment objects (see Section 2.2), represent them as persisting (see Section 3.2) and track their causal interactions. When and how do humans first track causal interactions?

In this section we will see that humans can do this from 4 months of age or earlier, and that the Simple View once again provides a candidate explanation of how they do it.

What sort of causal interactions might you need to track in order to know simple facts about particular physical objects? You probably do not need to be able to track particularly complex causal interactions. But it is plausible that you might need able to track very simple causal interactions, such as the collision of two balls or the interaction of a ball with a barrier.

Consider observing the scenario depicted in the middle of Figure 3.3. There is a bench above the ground. A screen appears, hiding most of the bench from view and then a ball drops down from above, moving behind the screen. Finally, the screen is lowered. Where do you expect to see the ball? Adults mostly expect to see it on the bench. After all, the ball cannot pass through the bench as both it and the bench are solid. Expecting the ball to be on the bench manifests an ability to track causal interactions: the bench’s stopping the ball is a causal interaction, so if you are unable to track causal interactions, there is no reason why you should expect the ball to be on the bench. Do infants also expect the bench to stop the ball? To find out, Spelke et al. (1992) first habituated 4-month-olds to a scenario with no bench. In this scenario, infants watch as a screen goes up, a ball falls behind it and then the screen comes down to reveal the ball on the ground (as depicted in the left part of Figure 3.3). Having been habituated to this scenario, infants were then shown a new scenario involving a bench. There were two versions of the new scenario. Some infants were shown a version of the new scenario in which the ball appeared on the bench, as adults would expect it to. This is depicted in the middle of Figure 3.3; call it the ‘consistent scenario’. Other infants were shown a version of the new scenario exactly like the other one except that when the screen came down, the ball appeared under the bench (as depicted in the right part of Figure 3.3); call this the ‘inconsistent’ scenario. To see whether 4-month-olds, like adults, expect the ball to be on the bench, Spelke et al. (1992) measured how much dishabituation each scenario produced. They reasoned that if infants were unable to track the ball’s causal interaction with the bench, then the consistent scenario should produce greater dishabituation because, when considered purely in terms of features, the consistent scenario is more different from the scenario to which infants were habituated than the inconsistent scenario is. Accordingly, if the inconsistent scenario produces greater dishabituation, that would be evidence that infants are not only representing causally inert features but can track the ball’s causal interaction with the bench. And this is just what they found: the inconsistent scenario was significantly more effective in arousing infants’ interest than the consistent scenario.

This is evidence that, by 4 months of age at the latest, infants can track simple causal interactions among objects, even when those causal interactions are occluded. Further evidence is provided by Baillargeon’s drawbridge study (Baillargeon 1987). I introduced this study in Section 3.2 as showing that infants can represent objects as persisting. But the study simultaneously supports the view that infants can track causal interactions: after all, it is not the mere presence of the object behind the drawbridge that matters but also its ability to constrain the drawbridge’s movement. (For further evidence, see Saxe, Tzelnic and Carey 2006.)

Figure 3.3 Scenarios used in an experiment on 4-month-olds’ abilities to track causal interactions. A ball falls behind a screen and then the screen is removed to reveal the ball in various positions.

Source: Spelke et al. (1992), Figure 2.

How do 4-month-old infants track causal interactions among objects? The Principles of Object Perception may be adequate to describing which causal interactions they detect. To illustrate, recall the principle of continuity. According to this principle, each object traces a connected path over space and time without crossing any other objects’ paths (Spelke et al. 1995, 113). The possible positions of a ball falling downwards when there is a bench in its path are limited by this principle. For the ball to move through the bench would involve a solidity violation of the principle of continuity, so the principle tells us that the ball cannot end up under the bench.

But to say that the principle of continuity and other principles describe how infants track causal interactions is not yet to explain how infants do this. The fact that certain principles can be used to describe infants’ behaviours does not force us to accept that those principles are in any sense guiding the behaviours. However, our current working hypothesis, the Simple View, commits us to taking the leap from description to explanation. According to the Simple View, the Principles of Object Perception (which include the principle of continuity) are things that infants know or believe, and infants can form beliefs about particular physical objects and their movements and interactions by using the principles in inferences. The Simple View therefore offers us a candidate explanation of how it is that infants (and adults) are able to track some causal interactions. But is the explanation correct?

3.5 The case for the Simple View

The main question for Part I of this book is: How do humans first come to know simple facts about particular physical objects, facts such as that this mug is over there? All the evidence we have considered so far points to an attractively uncomplicated answer in the form of the Simple View. According to this View, from 4 months of age or earlier, infants know a small number of principles about physical objects, their movements and interactions. And they are able to combine these principles with sensory information inferentially, thereby acquiring beliefs or knowledge about particular physical objects’ boundaries, locations and causal interactions.

Before considering a challenge to the Simple View in Chapter 4, let us pause to analyse the leaps of reasoning that take us to it. This is helpful both for evaluating the strength of the case for the Simple View and for working out what rejecting the Simple View would entail.

We started by asking a relatively simple question in Section 2.1. When and how do humans first manifest the three abilities associated with knowledge of objects? As we have seen, a variety of evidence indicates that humans manifest all three abilities by 4 months of age. Four-month-olds can segment objects, represent them as persisting and track some of their causal interactions. This is an extraordinary discovery. It used to be quite widely held that the infants lived in a world of mere features until much later in their development, and that these abilities were a consequence of learning through purposive interactions with physical objects (Flavell 1963). The grounds for that view have been almost entirely swept away by the newer research we have been considering. But the new discoveries leaves us with a question. Since it is not plausibly a consequence of learning through purposive interactions with physical objects, how do 4-month-old infants segment objects, represent them as persisting and track some of their causal interactions?

An important step towards answering this question is Spelke’s discovery that all three abilities in infants—to segment objects, represent them as persisting and track some of their causal interactions—can be described by invoking to a single set of principles, the Principles of Object Perception. (The principles discussed in this chapter are listed in Table 3.1.) This is a revolutionary discovery insofar as infants’ abilities to segment objects were previously thought to depend on information about shape and texture rather than being characterized by facts about physical objects’ nature, movements and causal interactions. This discovery also hints that infants’ abilities to segment objects, to represent them as persisting and to track their causal interactions may all be consequences of a single mechanism. As Carey and Spelke put it, ‘[a] single system … appears to underlie object perception and physical reasoning’ (1994, 175).

But of course we want more. We want to understand what kind of ‘system’ this is. To this end, it is useful to distinguish three claims about the Principles of Object Perception (summarized in Table 3.2). The first is that they are formally adequate. That is, someone who took the

Table 3.1 Some principles which partially characterize how infants segment physical objects, represent them as persisting, and track their causal interactions

PrincipleCharacterization
No action at a distance‘Objects are interpreted as moving independently of one another’
Rigidity‘Objects are interpreted as moving rigidly’
Cohesion‘Two surface points lie on the same object only if the points are linked by a path of connected surface points’
Boundedness‘Two surface points lie on distinct objects only if no path of connected surface points links them’
Continuity‘An object traces exactly one connected path over space and time’ and objects’ paths cannot cross

Source: Spelke (1990) and Spelke et al. (1995).

Table 3.2 Three questions about the Principles of Object Perception, if they are able to segment objects, represent them as persisting and track their causal interactions

Formal AdequacyIf someone took the principles to be true, was omniscient about the arrangement of surfaces in space and had unlimited cognitive resources, to what extent would she be able to segment physical objects, represent them as persisting and track their causal interactions?
Descriptive AdequacyDo the principles enable us to generate correct predictions about infants’ and others’ abilities to segment physical objects, represent them as persisting and track their causal interactions?
Explanatory AdequacyIs there a link between the principles and infants’ or others’ minds, and does this link partly explain how it is they are able to segment physical objects, represent them as persisting and track their causal interactions?

principles to be true, was omniscient about the arrangement of surfaces in space and had unlimited cognitive resources could, within limits at least, use the principles to segment objects, represent them as persisting even while briefly hidden from view and track their causal interactions. While it is unlikely that anyone has yet formulated principles that are formally adequate in this sense, the principles we do have suffice to give us a sense of what a formally adequate set of principles would look like. To say that the Principles of Object Perception are formally adequate is not yet to say anything at all about how these principles relate to infants’ abilities.

A further claim is that the principles are descriptively adequate for capturing infants’ and others’ abilities. That is, these principles allow us to generate correct predictions about how infants, adults and many nonhumans will interpret particular scenarios. They tell us, for instance, how many moving objects someone will interpret a scenario as containing, and where she will locate those objects at different times. If, as I assume, the principles are not too far from being formally and descriptively adequate, then we could in principle use them to build a machine whose abilities regarding physical objects matched those of infants. But in building such a machine we would not necessarily be replicating the inner workings of an infant. The claims about formal and descriptive adequacy are not candidate answers to questions about what underlies infants’ abilities. That certain principles are descriptively adequate for capturing infants’ abilities to segment objects, represent them as persisting and track some of their causal interactions does not yet tell us how it is that infants are able to do these things.

This is where the third claim comes in. To say that the principles are explanatorily adequate for capturing infants’ or others’ abilities is to say that there is a link between the principles and their minds, and that it is partly in virtue of this link that they have these abilities. Where such a link exists, the principles do not merely describe the infants’ or others’ abilities: they explain them by virtue of their role in characterizing processes, representations or systems underlying the abilities. While researchers disagree on how the principles are linked to infants’ minds (as we will see in Chapter 4), many would probably accept the broad claim that the Principles of Object Perception are explanatorily adequate. We too should accept this claim, at least provisionally, not because the available evidence overwhelmingly supports it but rather because there is currently no better supported alternative.

The claim that the Principles of Object Perception are explanatorily adequate is too schematic to be satisfying. Accepting this claim requires us to ask, What links the principles and infants’ minds in such a way that infants are able to segment objects, to represent them as persisting and to track their causal interactions?

The Simple View is one answer to this question. It is not the only possible answer: accepting that the Principles of Object Perception are explanatorily adequate does not compel us to accept the Simple View. But the Simple View is the only way of answering the question we have yet considered. On the Simple View, the relation between the principles and infants’ minds is one of knowing or believing. It is in virtue of infants’ knowledge of, or belief in, the Principles of Object Perception that they are able to segment objects, represent them as persisting and track their causal interactions.

As I keep saying, the Simple View provides an answer to our main question, How do humans first come to know facts about particular physical objects, such as their locations? According to the Simple View, they do this by a process of inference which combines principles they know or believe with sensory information about arrangements of shapes in space. This is an attractive answer insofar as it postulates only states and processes that are familiar and that are needed to explain many other things. It also has the advantage of entailing an uncomplicated thesis about how infants’ earliest abilities relate to adults’ abilities: they are identical. Unfortunately, as we are about to see, the Simple View is unlikely to be true.

Notes

4 The Linking Problem

How do humans first come to know simple facts about particular physical objects? Knowing any such facts probably involves being able to segment physical objects, to represent them as persisting, and to track their causal interactions. In Chapter 2, we saw that infants can do all of these things from around 4 months of age or earlier. We also saw that there is a single set of principles, the Principles of Object Perception, which describe how infants do these things. This led us to ask, What links these principles to infants’ minds? Our current answer is the Simple View. According to the Simple View, infants know or believe the Principles of Object Perception. And they use these principles in inferring facts about physical objects, their locations, movements and interactions, from sensory information about features. The Simple View promises an attractively straightforward answer to the question about how humans first come to know facts about particular physical objects: they do so in a way that they might later come to know facts about economic values or legal obligations, namely, by inference.

This chapter is about why the Simple View is wrong, and about the problem we face when we reject the Simple View, which I’ll call the Linking Problem. The Simple View is wrong because it makes incorrect predictions (see Sections 4.1–4.3). The problem arising from its failure, the Linking Problem, is to provide an alternative account of what links the Principles of Object Perception to infants’ and others’ minds and so explains their abilities to segment objects, represent them as persisting and track their casual interactions (Section 4.4). If not belief or knowledge states, what does link the principles to particular minds?

Devoting a whole chapter to explaining a problem might seem excessive, even given that some theoretical foundations will be laid along the way. But versions of the Linking Problem arise for all domains of knowledge. It is a pervasive but often overlooked obstacle to understanding how knowledge emerges in humans.

4.1 Against the Simple View

Some philosophers have used intuitions to develop sophisticated arguments against the Simple View. Bermúdez (2003), for instance, argues that those without the ability to use a language cannot make inferences; and Davidson argues that those without language cannot think at all.201 It may be hard to accept that 4-month-old infants are in the business of inferring truths about particular objects’ locations from abstract principles. (And perhaps it is no less hard to accept that adults typically do this in segmenting objects.) But scientific and mathematical discoveries sometimes require us to reject intuitions, perhaps even deeply held intuitions about very fundamental things like space and time. For this reason there seems to be slim prospect of effectively challenging the Simple View on the basis, ultimately, of intuitions about the nature of knowledge, belief and inference. In any case, doing so is also unnecessary. For there are excellent scientific reasons for rejecting the Simple View.

Recall Baillargeon’s (1987) experiment which used a rotating drawbridge to show that 4-month-olds can represent objects as persisting even while briefly hidden from their view. (The scenarios used in this experiment are depicted in Figure 3.2.) As we saw, converging evidence comes from other experiments using different scenarios and different means of detecting infants’ responses to them. The means include dishabituation, looking times as a marker of violation-of-expectation, and anticipatory looking; there is also a study involving neural measures (Kaufman, Csibra and Johnson 2005; this study will concern us in Chapter 6). But what happens if instead of measuring infants’ looking or neural responses, we instead measure how they search for objects?

Shinskey and Munakata (2001) did just this. They created apparatus much like that used by Baillargeon (1987) in her drawbridge studies. There was an opaque screen that could rotate between lying flat on the ground and being raised to conceal either a toy (in some conditions) or nothing (in other conditions) behind it. Stops prevented the screen from ever moving into a flat position even if there was nothing behind it. So the position of the screen alone would never reveal whether there was a toy behind it. The key difference between this study and Baillargeon’s was that the infants could choose to pull the screen forwards towards themselves, thereby revealing whatever is behind it. Shinskey and Munakata (2001) also used a second piece of apparatus just like the first except that the screen was transparent rather than opaque. They reasoned that infants would quite often pull the screen forwards just for fun, regardless of what is behind it. However, they also guessed that when infants know there is an interesting toy behind the screen, then they would pull it forwards more often than when they knew that there is nothing behind the screen. This is just what happened when infants were presented with the apparatus involving a transparent screen: they sometimes pulled the screen forwards when there was no toy behind it, but they pulled it forwards significantly more often when the toy was behind it (see the left bars, labelled ‘transparent’, in Figure 4.1). What happened when infants were presented with the opaque screen? Here infants pulled the screen forwards no more often when they had observed a toy being placed behind it then when they had observed that there was nothing behind it (see the right bars, labelled ‘opaque’, in Figure 4.2). This is evidence that 7-month-old infants do not know

Figure 4.1 The transparent and opaque barrier apparatus used in an experiment showing that 7-month-olds will not search for an object after it is hidden from their view

Source: Shinskey and Munakata (2001), Figure 1.

Figure 4.2 Seven-month-old infants’ actions indicate that they are unaware of a toy after it is hidden behind an opaque screen. Note that what matters is not the absolute percentage of screen pulls but whether there is significant increase in screen pulls when the toy is present behind the screen. When the screen is opaque, there is no such increase.

Source: Shinskey and Munakata (2001), Figure 2, part.

that a toy they have very recently seen hidden behind a screen is behind the screen. After all, since knowledge guides action, we would expect infants who know that a toy is behind an opaque screen to pull the screen forwards more often than infants who know there is nothing behind the screen, just as they do when the screen is transparent.

This is a problem for the Simple View. The Simple View predicts, correctly, that infants’ looking behaviours will reveal sensitivity to the locations of physical objects that are momentarily out of view. But the Simple View also predicts, incorrectly, that the same infants’ searching behaviours will likewise reveal sensitivity to the locations of physical objects. In fact, even much older infants, 7-month-olds, do not search for hidden objects. The Simple View generates an incorrect prediction.

Our knowledge that infants fail to act on objects hidden behind a barrier or a screen does not rest on just one study. I described Shinskey and Munakata’s (2001) experiment because it is so elegant, but more than two decades of research strongly supports the view that infants fail to search for objects hidden behind impenetrable barriers or screens until around 8 months of age (Meltzoff and Moore 1998, 202) or maybe even later (Moore and Meltzoff 2008).

Someone might attempt to rescue the Simple View by postulating factors which would prevent infants’ knowledge from manifesting itself in their searching for the hidden object. The idea is that infants do know that the toy is behind the opaque screen but they are prevented from searching for it by some extraneous factor. It is difficult to identify any such extraneous factor, however. Could it be that infants are unable to prepare or perform the searching action? Shinskey and Munakata (2001) rule out this possibility by including the transparent screen apparatus. Since both opaque and apparent apparatus demand the same actions, the difference in infants’ behaviours cannot be due to difficulties in acting. As Shinskey argues on the basis of some further experiments, ‘action demands are not the only cause of failures on occlusion tasks’ (2012, 291). Further evidence that infants’ failure to search for a hidden object are not due to difficulties preparing or performing a search action is provided by Moore and Meltzoff (2008), who compare infants’ abilities to search for partially hidden and fully hidden objects.

Could it be that infants do know the toy is behind the opaque screen but are prevented from reaching for it by the twin of demands of having to hold their knowledge in mind and to prepare an action? For comparison, suppose you are asked to hold an eight digit number in mind just a moment after your phone has rung. Even if you know where the phone is, you might be briefly prevented from reaching for it by the sheer difficulty of holding the digits in mind. Could something similar be true of infants? Could the non-visibility of the object place demands on memory that consume the cognitive resources also required for preparing action? Moore and Meltzoff (2008) reasoned that if this were right, having the toy make a noise continuously should help infants to remember it and so boost their performance. But they found that 8-month-olds failed to search for a toy irrespective of whether it made a noise or not (Moore and Meltzoff 2008, Experiment 2). This suggests that memory is unlikely to be the only limiting factor.

There is also evidence that, when an object is not hidden by an impenetrable screen, infants even younger than the 7-month-olds Shinskey and Munakata (2001) tested can cope with the twin of demands of having to hold their knowledge in mind and to prepare an action.

To demonstrate this, Shinskey (2012) hid objects in a tray of milk. She found that 6-month-olds reach into the milk significantly longer after observing that a toy is hidden in the milk than after not observing this. Relatedly, Jonsson and Von Hofsten (2003) compared infants’ reaching to an object in darkness with their reaching to an object momentarily hidden by an impenetrable barrier. They found that 6-month-olds will not reach for an object hidden by a barrier but will reach for one hidden by darkness (see also Hespos et al. 2009; Babinsky, Braddick and Atkinson 2011). The fact that 7-month-olds can, in some conditions, reach for an object they are not currently perceiving indicates that demands on holding knowledge of an unperceived object’s location in mind do not explain why they fail to retrieve objects from behind a barrier. It seems, then, that infants’ failures to act really do indicate that they do not know where an object hidden behind a barrier is. And this means we must reject the Simple View.

To recap, according to the Simple View, 4-month-olds know where an object recently hidden behind an opaque screen or barrier is, by 4 months of age at the latest. This view correctly predicts infants’ responses to the presence of an object behind an opaque screen in studies involving habituation, violation-of-expectation and anticipatory looking methods. But it also predicts that the same infants will search for an object recently hidden behind an opaque barrier. This prediction is false. Infants will not do this until around 8 months of age at the earliest. We could rescue the Simple View if we could identify an extraneous factor which prevented infants from manifesting their knowledge in searching for objects hidden behind an opaque barrier. But identifying any such extraneous factor is difficult because infants have no difficulty in searching for an object behind a barrier when the barrier is transparent, and they have no difficulty searching for an object when it is hidden by milk or darkness rather than by an impenetrable barrier. Despite much effort, no one has published a convincing explanation of infants’ failures to search that is consistent with the claim that they know where an object recently hidden behind an opaque screen is. For this reason, we should reject the Simple View.

Or should we? Even without being able to explain away the Simple View’s failure, you might refuse to abandon it on the basis of just one incorrect prediction.

4.2 Further evidence against the Simple View

We have just seen that the Simple View makes an incorrect prediction about infants’ abilities to represent objects as persisting. The Simple View also makes an incorrect prediction about infants’ abilities to track causal interactions, or so this section will argue. Taken together, the two incorrect predictions provide strong grounds for rejecting the Simple View.

Consider a scenario in which a ball rolls down a ramp. The ball is stopped by a barrier, which can be placed in various positions along the ramp. In front of most of the ramp and barrier there is a screen with four doors in it. The screen is tall enough to hide the last part of the ball’s journey, but the barrier sticks up over the screen making it easy to predict where the ball will stop—at least, this is easy for adults. The barrier is always placed after one of the four doors in the screen, as depicted in Figure 4.3. Where will infants expect the ball to stop?

Figure 4.3 Figure 4.3 A ball rolls down a ramp and hits a barrier. Which door will it end up behind?

Source: Hood, Cole-Davies, and Dias (2003), Figure 1.

If you recall Spelke et al.'s (1992) experiment in which a ball drops vertically behind a screen with a solid bench in its path (as depicted in Figure 3.3), you might guess that even 4-month-old infants will correctly expect the ball to stop in front of the barrier. This is probably half-right. In one experiment Hood, Cole-Davies and Dias (2003) showed 2- and 3-year-olds the ball rolling down the ramp. They then opened a door to reveal the ball. Some of the time they contrived to retrieve the ball from the wrong door, as if by magic. In these cases the children looked significantly longer than when the ball was retrieved from the correct door. This indicates that these children expected the barrier to stop the ball. Mash et al. (2006 Experiment 1) refined Hood, Cole-Davies and Dias's (2003) procedure. Rather than sometimes retrieving the ball from an incorrect location, they simply opened doors. But they contrived things in such a way that sometimes the ball was not behind the correct door. They found that when the correct door was opened, infants looked longer when no ball was behind it than when the ball was behind it. This is evidence that infants anticipate the ball's location. It is not just that seeing the ball in the wrong place enables them to realize something is wrong: merely seeing the ball's absence is enough to confound their expectations.

So far these findings accord with those discussed earlier (even if the subjects are 2 or 3 years old rather than 4 months), and with the predictions of the Simple View. But the useful feature of the rolling ball scenario is that it allows us to investigate infants' expectations using two different measures. What happens if we test where 2-year-olds expect the ball to stop not by measuring looking times but instead by asking them to open a door to retrieve the ball?

When Berthier et al. (2000) did this, they found that 2-year-olds did not open the correct door more often than someone opening doors at random would. Some children selected a favourite door which they always opened; others always opened a door adjacent to the barrier but had no preference for the door on the correct side of the barrier. This failure to search for the ball in the correct location is evidence that 2-year-olds do not expect the solid barrier to stop the moving ball and do not know where the ball is at the end of its journey.

Because this is such an unexpected finding, Berthier et al.'s (2000) experiments have been replicated and extended by several labs. Hood, Cole-Davies and Dias (2003) gave the same individuals both violation-of-expectation and search tasks using the same apparatus. All children looked longer at apparently impossible events in which a ball was retrieved from a wrong location, but few of their 2½-year-olds reliably searched for the ball in the correct location. Butler, Berthier and Clifton (2002) modified the screen in front of the ramp so that it was entirely transparent except for the doors. This meant that children could observe the ball rolling between the doors. Amazingly, this had only a modest effect on 2-year-olds' performance: most (70 per cent) still did not reliably search for the ball in the correct location. Keen et al. (2008) explored three different ways of getting children to focus on the barrier, reasoning that younger children might fail the search task simply because they were not paying attention to the barrier. It turned out that none of these much improved the children's performance in searching for the ball. The truth is that while 4-month-olds' looking behaviours indicate sensitivity to the fact that solid barriers stop moving objects, 2-year-olds' searching actions indicate ignorance of this fact.

The same discrepancy between looking and searching as evidence for abilities to track causal interactions has been found in adult nonhuman primates, specifically cotton-top tamarins (Santos, Seelig and Hauser 2006). Related discrepancies have also been found in other adult nonhuman primates (Gómez 2005; Santos and Hood 2009).202 This suggests that discrepancies between looking and search behaviours are not a consequence of extraneous factors or capacity limits. After all, such discrepancies should never occur in fully developed adults of any species if they were. Discrepancies between looking and search behaviours in some adult primates therefore strengthen the case against the Simple View.

4.3 Things get even worse for the Simple View

All the evidence against the Simple View considered so far arises from patterns in infants' abilities to perform purposive actions. By contrast, much of the evidence in favour of the Simple View comes from paradigms which involve eye movements that are either very simple purposive actions or not purposive actions at all. This invites a possible defence of the Simple View. Could tasks requiring non-ocular purposive actions be less sensitive measures of infants' capacities? If so, the Simple View might still be broadly supported by the available evidence.

The problem with this line of defence is that infants' abilities are sometimes manifest in tasks requiring reaching but not in looking time measures. As mentioned earlier (in Section 4.1), infants will reach for an object hidden in darkness; given the Simple View, this indicates that they are able to know the object's location even while it is hidden. But what happens if instead of measuring reaching, we measure looking times? Charles and Rivera (2009) compared what happens when an object is momentarily hidden behind a screen with what happens when an object is momentarily hidden by darkness. They used a trick with light and mirrors so that for some of the infants, the object did not reappear when the screen came up or the light returned. Surprisingly, 5-month-old infants' looking times indicated that an expectation had been violated only when the object was hidden behind a screen and not when it was hidden by darkness.

Apparently, then, tasks requiring eye movements are not more sensitive than tasks which require manual actions. Depending on the scenario used, infants in their first five or so months of life will fail to reveal capacities to track physical objects in either kind of task.

The Simple View therefore generates incorrect predictions not only about manual search behaviours but also about performance on violation/of/expectation tasks, as Table 4.1 summarizes. If, as infants' searching behaviours suggest, they know roughly where the object hidden by darkness is, why is this knowledge not manifest in their looking times? And if, as infants' looking behaviours indicate, they know roughly where an object occluded by an impenetrable barrier is, why is this not manifest in their manual actions?

The case against the Simple View, although not decisive, is compelling. According to the Simple View, humans know or believe principles which partially characterize the boundaries, movements and interactions of physical objects, and they are able to use these principles to infer truths about the locations of objects. This view does make correct predictions in some cases, but it also makes systematically incorrect predictions about infants' and children's purposive actions. If the Simple View were correct, 7-month-olds should be able to retrieve objects from behind an opaque barrier, and they should be able to correctly locate a ball whose path is manifestly blocked by a solid barrier. They should also look longer when a briefly endarkened object fails to reappear. That they cannot do any of these things is a compelling reason to reject the Simple View.

Infants in the first months of life do manifest symptoms associated with knowledge about particular physical objects. But this is not the same thing as their actually knowing things (Section 1.2).

Rejecting the Simple View leaves us with a cluster of questions. How else can we explain 4-month-olds' abilities to segment objects, represent them as persisting and track their causal interactions? Why do infants manifest these abilities when they are measured in some ways but not others? How else are the Principles of Object Perception linked to infants' minds? And since it is not by inference from principles known or believed, how else is it that humans first come to know simple facts about particular physical objects? These are all aspects of the Linking Problem.

4.4 The Linking Problem

Infants' abilities to segment physical objects, to represent them as persisting even while briefly unperceived, and to track their causal interactions can be described by the Principles of Object Perception (see Section 2.3). But what links these principles to infants' minds in such a way as to explain these abilities?

Consider the three most familiar kinds of mental representation, which are perceptual, motoric and epistemic. (These were introduced in Section 1.3.) Perceptual representations are those involved in perceiving. It may seem obvious, and it is often more or less taken for granted, that infants' abilities to represent physical objects as persisting even while briefly unperceived cannot be explained by appeal to perceptual states. After all, much of the point of considering briefly unperceived objects is to rule out the possibility that infants' abilities are merely perceptual.

What about motor representations? Could these explain infants' abilities concerning physical objects? Motor representations are the states involved in preparing and performing very small-scale actions like reaching for a cup, grasping it, transporting it to your mouth and drinking from it (see Sections 1.3 and 11.2, for more on motor representations). Objects are indeed sometimes represented motorically (see Section 6.7), but usually not when they are unavailable to act on, as, for instance, when they are behind an impenetrable barrier. This makes it unlikely that 4-month-olds' abilities concerning physical objects could be entirely a consequence of motor representations.

Eliminating perceptual and motor representations leaves us with the third familiar kind of mental representation: epistemic. This includes things like knowledge, belief, desire and intention. The Simple View is the view that 4-month-old infants' abilities are underpinned by epistemic states. We know this is probably false because, as we have just seen, the Simple View generates multiple incorrect predictions.

Our question is, What links the Principles of Object Perception to infants' (and perhaps others') minds? Identifying the link is essential for explaining how they are able to segment physical objects, represent them as persisting and track their causal interactions. And it is a problem—call it the Linking Problem—because it seems we cannot identify the link by invoking any of the three familiar kinds of mental representations.

Davidson might be interpreted (or constructively misinterpreted) as describing the Linking Problem in a passage reflecting on philosophical challenges posed by development:

The difficulty in describing the emergence of mental phenomena is a conceptual problem: it is the difficulty of describing the early stages ... that precede the situation in which concepts like intention, belief, and desire have clear application ...

We have many vocabularies for describing nature when we regard it as mindless, and we have a mentalistic vocabulary for describing thought and intentional action; what we lack is a way of describing what is in between.

(Davidson 2001, 127–8; compare Davidson 1999b, 11)

The Linking Problem arises in attempting to explain the emergence of knowledge, which is a mental phenomenon. The failure of the Simple View indicates that Davidson is right insofar as solving the Linking Problem requires a way of identifying something that is 'in between' mindless behaviour and epistemic states such as knowledge or belief. But Davidson takes it to be a 'conceptual problem'. And he takes the problem to be unsolved, perhaps even unsolvable. In what follows we will see grounds for thinking that the Linking Problem is not a narrowly conceptual problem as it can be solved, in at least one domain, by careful attention to scientific discoveries (see Chapter 6).

4.5 Representation not knowledge

It may be tempting to think that the Linking Problem is trivial or already solved. Don't scientists already have a perfectly good vocabulary for describing what is in between mindless behaviour and knowledge or belief? Isn't talk about representation supposed to serve just this purpose?

Thinking along these lines, we might aim to retain many of the virtues of the Simple View by switching from talking about knowledge and belief to talking about representation. Instead of saying, as the Simple View does, that the Principles of Object Perception are things 4-month-olds know or believe, we can say instead that they represent these principles. And, relatedly, we can say that these principles feature in inference-like cognitive processes that lead to representations of facts about particular physical objects, such as facts about their boundaries, locations and interactions. Since not all kinds of representation are linked to purposive action in the way that knowledge and belief are, this more cautious view does not generate the false predictions that the Simple View generates.

Invoking representation instead of knowledge or belief is a good move insofar as it enables us to avoid a view that clearly generates incorrect predictions. But it is not a move that, all by itself, will solve the problems arising from the failure of the Simple View. This is because invoking representation instead of knowledge or belief merely amounts to switching from a more concrete view that provides candidate explanations and generates testable predictions to a less concrete view that explains little and predicts less.

Representation is a generic notion. Photographs, maps, blueprints and sentences are all representations, and all have mental counterparts. If you wish, want, hope, intend, imagine, guess, fear, believe or know that Sam will have an easy birth, then you represent this. Little can be predicted about your behaviour from the bare fact of your representing that Sam will have an early birth. Haith (1998) claims that 'no concept causes more problems in discussions of infant cognition than that of representation'. The first step towards avoiding these problems is to recognize how little can be predicted or explained with the bare fact that someone represents something.

Imagine observing as Ayesha knocks her would-be mugger out cold with a single, well-aimed swing of her handbag. Guessing that Ayesha has loaded her handbag with rocks, you pick it up yourself but discover that it feels much too light to contain rocks. You must change your guess about what is in her handbag (unless, of course, you are willing to postulate some extraneous factor interfering with the normal effects of gravity on rock). Changing your guess from rocks to physical matter is a step forward insofar as you are no longer obviously wrong. But it is also a step backward insofar as your new guess does not provide an interesting candidate explanation for the knock-out blow. It is barely controversial that Ayesha's handbag contains something physical. Of interest is which kind of physical thing has this combination of low density and clout. Representation is a generic notion in the way that physical matter is.

Where psychologists tend to use 'representation', philosophers are perhaps more likely to talk about 'tacit' or 'implicit' knowledge. Despite much research on this topic stretching across decades (Stich 1978 is a particularly brilliant, pioneering example) and some sustained attempts taking a variety of different approaches (including Davies 1989; Dienes and Perner 1999; Sperber 1997, among many others), it appears that we have all made quite limited progress in characterizing such notions. My own sense is that the most we can safely say is that tacit or implicit knowledge is like knowledge but lacks some of the features associated with it. If this is our position, switching from a claim about knowledge to one about tacit or implicit knowledge is like switching from knowledge to representation. These switches are ways of marking our ignorance about what links the Principles of Object Perception to the minds of individuals. They do not solve the Linking Problem.

What kind of representations (if any) do 4-month-old infants have of physical objects? What is it about this kind of representation that enables them to give the appearance of knowing things about physical objects on some tasks even while failing other tasks which, apparently, test for the same knowledge (see Table 4.1). What is the relation between the representations of this kind, which infants have by at least 4 months of age, and the knowledge of physical objects which they lack until months or years later? Until we can answer these questions, we have not solved the Linking Problem.

4.6 Graded representations?

Faced with the difficulty of solving the Linking Problem, we might reasonably attempt to duck it. Does explaining 4-month-olds' abilities to segment objects, represent them as persisting and track their causal interactions really require any kind of mental representation that is not knowledge?

A radical way to duck Linking Problem would be to attempt to explain 4-month-old infants' abilities without relying at all on conjectures about knowledge or mental representations. This idea is outlined by Smith (2005) and Schöner and Dineva (2007). As these authors admit, the approach they are aiming to elaborate is not currently well understood and their theories are at a comparatively early stage of development. Appeals to mental representation may eventually turn out to be as mistaken as invoking vital forces or ether. But while there is much intriguing theoretical and experimental research covering isolated phenomena, at present, there is probably no good alternative to starting with mental representations if our aim is to understand how knowledge emerges in development.

A different way to duck the Linking Problem has been proposed by Munakata (2001; see also Munakata et al. 1997). She suggests that knowledge can be 'graded': some knowledge states are 'stronger' while others are 'weaker'. She also holds that weaker knowledge states can drive looking time behaviours but not control purposive action. This allows her to hold that 4-month-olds know principles about objects generally and facts about particular objects, just as adults do. On this view, infants' representations of objects are not different in kind from adults' knowledge of facts about particular physical objects.

The idea that knowledge can be graded is initially attractive. It appears to avoid the incorrect predictions of the Simple View, yet provides a simple solution to the Linking Problem. But this idea also faces a challenge. Talk about 'strength' in this context needs to be anchored in a theory of representation. When talking about a radio signal, say, it is possible to specify what signal strength amounts to, and it is clear that signals of varying strength can all carry the same message. But what does strength amount to in the case of a mental representation or knowledge state?

Without an answer to this question, invoking graded representations will not explain anything. It amounts merely to retrospectively postulating a novel aspect of representation to characterize findings about what 4-month-olds can and cannot do.

To see the force of this challenge, consider that proponents of graded knowledge hold that weaker representations can guide many looking behaviours, whereas manual search behaviours generally require stronger representations. Why is it this way around? Why is a stronger representation needed for manually searching than for looking? And why is a stronger representation needed when objects are occluded rather than merely endarkened? Until we can answer such questions, postulating graded knowledge will not explain the developmental puzzles about knowledge of objects.

Another problem is that infants at around 4 or 5 months of age do not always fail to manually search for objects, nor do they always succeed in manifesting an appearance of knowledge in their looking behaviours (see Table 4.1). In solving the Linking Problem. it is not enough to explain why infants sometimes fail to search. We must also explain why they sometimes succeed in searching yet fail in looking.

One approach to explicating what strength is might be to equate it with confidence. What is confidence? To illustrate, consider Ayesha who knows she has never visited Milan and also that she has never been to the moon. While she knows both things, she is much more confident about the second. Her confidence is reflected in the fact that she would risk more if offered an opportunity to bet on whether she has ever been to the moon. Can we equate strength with confidence? This is not the view of any proponent of graded representations, but since we already know that it is necessary to postulate degrees of confidence, equating strength with confidence would avoid introducing a new aspect of representation.

Unfortunately equating strength with confidence is unlikely to work. This is because differences in confidence are unlikely to explain why infants' representations of momentarily hidden objects influence many of their looking behaviours but have no effect on many of their manual searching behaviours. Consider, for instance, a 2-year-old infant who can open one of several doors and will get a reward if she opens the door with an object behind it (as in the paradigm of Berthier et al. (2000) discussed in Section 4.2). If she knows which door the ball is behind, then she should not be systematically opening the wrong door, however low her confidence might be. So equating strength with confidence would mean that invoking graded representations generated incorrect predictions.

How else might the notion that representations vary in strength be explicated? An ambitious move would be to anchor the notion that representations can be stronger or weaker in a connectionist model of infants' performance on some tasks involving briefly hidden objects (see Munakata et al. 1997, 698), or in facts about the neural basis of representations (see Munakata 2001, 309). Success in doing this has the potential to provide a robust theory of graded representations, one likely to generate many distinctive predictions. Unfortunately there are several challenges facing this approach. The first is to explain which patterns in neurons' firing determine strength. The second challenge is to explain why, when strength is understood in terms of patterns in neurons' firing, more strength should generally be necessary for manual search behaviours than anticipatory looking. And the third challenge is to explain which feature of representations corresponds to the patterns in neurons' firing. Currently none of these challenges have been addressed in detail.

As things stand, appeal to the notion that knowledge states can vary in strength does not enable us to duck the Linking Problem. Nor is there yet any substantial prospect of explaining infants' abilities without appeal to mental representations of some kind. Of course the challenges facing approaches to ducking the Linking Problem might eventually be overcome. But for now our best hope of understanding how knowledge of simple facts about physical objects emerges in development is probably to face the Linking Problem head on.

4.7 Conclusion

In order to understand how humans first come to know simple facts about particular physical objects, we need to understand infants' abilities to segment objects, represent them as persisting and track their causal interactions. An initial attempt, the Simple View, generates incorrect predictions. The problem with the Simple View is that it entails that infants who can segment objects, represent them as persisting and track their causal interactions thereby also know facts about the locations, movements and interactions of particular physical objects. But there is considerable evidence that they do not. The former abilities can exist in the absence of any corresponding knowledge about physical objects, as we saw in Sections 3.1 and 3.2.

The failure of the Simple View means we are confronted with the Linking Problem. We have to identify a link between principles that describe infants' abilities concerning physical objects and their minds. How can we do this?

A first idea was to switch from knowledge to representation: instead of saying that infants know or believe things about particular physical objects, we might say merely that they represent them. This probably enables us to avoid falsehood, but it won't enable us to explain the origins of knowledge. We need more (see Section 4.5). What is the nature of infants' earliest representations of physical objects? In particular, what is it about their representations in virtue of which they appear to manifest knowledge, when tested in some ways, while appearing to lack knowledge ,when tested in other ways? And what is the relation between these early representations and the knowledge of facts about particular physical objects which appears later in development?

It is just conceivable that we can duck the Linking Problem by invoking the idea that knowledge can have different grades of strength and weakness. But, as we saw in Section 4.6, the challenges that would have to be overcome in providing an explanatory account of graded representations capable of generating useful predictions are probably at least as daunting as the challenges involved in facing the Linking Problem head on.

What next? The most influential, best developed attempt to solve the Linking Problem involves the notion of core knowledge.

Notes

5 Core knowledge

In attempting to understanding how humans first come to know simple facts about particular physical objects, one important discovery is that 4-month-olds can segment objects, represent them as persisting and track their causal interactions. Further, their abilities to do these things can be described, at least approximately, by a set of principles about the behaviour of physical objects, Namely, the Principles of Object Perception (see Chapters 2 and 3). Given this, it is natural to wonder whether the Principles might not merely describe but also explain infants' (and adults') abilities (see Table 3.2).

If the Principles of Object Perception somehow explain infants' abilities, there must be some link between the Principles and individual minds. A natural first idea is that this link involves belief or knowledge: the Principles are things that infants believe or know. Add the further claim that infants acquire beliefs or knowledge about particular physical objects and their interactions by making inferences from these principles together with perceptual information about the arrangement of surfaces in space and we arrive at the Simple View (see Section 3.1).

The Simple View is attractive insofar as it does not require postulating novel kinds of mental states or processes. But it systematically generates incorrect predictions (see Chapter 4), so must be rejected. It appears we cannot understand how the principles might be linked to individual minds by any of the most familiar kinds of mental states—the perceptual, motoric and epistemic. The question of what links the Principles to individual minds is therefore a problem. This is the Linking Problem, introduced in Chapter 4. In this chapter we will investigate a leading attempt to solve it by postulating something called 'core knowledge'.

5.1 What is core knowledge?

In attempting to solve the Linking Problem, we have so far considered only the most familiar kinds of mental representations—knowledge states, beliefs, perceptions and the like (see Section 4.4). Does understanding the developmental emergence of knowledge require postulating a novel kind of mental representation?

Carey suggests it does. According to her: 'there is a third203 type of conceptual structure, dubbed “core knowledge” ... that differs systematically from both sensory/perceptual representation ... and ... knowledge' (2009, 10). Could core knowledge be what links the principles describing infants' earliest abilities concerning physical objects to their minds?

To answer this question, we first need to know what core knowledge is. Core knowledge is defined in terms of core systems. For someone to have core knowledge of a particular principle or fact is for her to have a core system where either the core system includes a representation of that principle or else the principle plays a special role in describing the core system.204 But what is a core system?

Start with an analogy: 'Just as humans are endowed with multiple, specialized perceptual systems, so we are endowed with multiple systems for representing and reasoning about entities of different kinds' (Carey and Spelke 1996, 517). As examples of 'specialized perceptual systems', Carey and Spelke mention things like perceiving colour, depth or melodies. The operations of such systems are to a significant extent independent of each other, and of capacities to know things about colour, depth or melodies. So the idea is that core systems are independent in whatever ways some perceptual systems are.

The analogy with perceptual systems is helpful for getting a rough, intuitive fix on what a core system is supposed to be. But what is a core system?

Core systems are standardly identified by listing their characteristic features. Carey and Spelke (1996, 520) write that core systems are:

1 largely innate 2 encapsulated 3 unchanging 4 arising from phylogenetically old systems 5 built upon the output of innate perceptual analysers.

The exact list of properties has varied over time, but not in ways that matter for our purposes. What matters for now is just that core knowledge is defined in terms of core systems, and core systems are defined by a list of features which probably includes most of those just mentioned.

In addition, core knowledge is sometimes held to be specific to quite narrow categories of event and not to grow by means of generalization (Baillargeon 2001, 344, 2002, 57). It may also be also best understood as a collection of rules rather than a coherent theory (Baillargeon 2002, 82). These claims do not follow directly from definitions: they are based on observations of limits on the abilities which, according to the theory, are underpinned by core knowledge.

Are the notions of core knowledge and core system useful tools for solving the Linking Problem? And will they enable us to describe—and maybe explain—what is in between mindless behaviour and more familiar kinds of mental states like knowledge, belief and perception?

Consider what I will call the Core Knowledge View:

The Principles of Object Perception are not things infants know but rather are encoded in their core knowledge. The operations of the core system enables this core knowledge, together with perceptual inputs, to generate expectations concerning particular physical objects. These expectations are not knowledge states but representations in core systems.

The Core Knowledge View is an alternative to the Simple View. According to the Simple View, infants know the Principles of Object Perception. But on the Core Knowledge View, infants may lack beliefs about, or knowledge of, simple facts about particular physical objects while having representations of such facts as a consequence of the operation of a core system. This is an advantage of the Core Knowledge View. It means that the Core Knowledge View will not generate the incorrect predictions which follow from the Simple View.

Not all researchers sympathetic to the existence of core systems would accept the Core Knowledge View. They may instead hold that the operations of core systems result directly in beliefs, or knowledge states, concerning particular physical objects (compare Leslie 1988, 193–4, on the outputs of modules). But holding this means that invoking core knowledge will not enable you to respond to the objections to the Simple View we have encountered in Sections 4.1 and 4.2). From our point of view, this would defeat the point of invoking core knowledge, which is to solve the Linking Problem.

5.2 Can core knowledge solve the Linking Problem?

Is the Core Knowledge View a solution to the Linking Problem? As we have just seen, the Core Knowledge View does not generate the incorrect predictions generated by the Simple View. Instead it has a complementary defect.

In evaluating the Simple View earlier in this chapter, we identified a puzzling pattern of evidence. The pattern is summarized in Table 4.1. A solution to the Linking Problem should enable us to explain this puzzling pattern. But consider the list of features used to define core systems in Section 5.1; they are largely innate, encapsulated, and so on. Which of these features explain the discrepancy between measures on which infants do, and measures on which they do not, manifest their abilities to track physical objects? Why do they fail on some manual search tasks and but pass some violation-of-expectation tasks when the mode of disappearance is occlusion? And, equally pressingly, why do they do the converse (pass search but fail violation-of-expectation tasks) when the mode of disappearance is endarkening (see Table 4.1)?

None of the features specified in defining core systems and core knowledge are relevant. The Core Knowledge View is consistent with the puzzling pattern of findings about infants' abilities concerning physical objects. But it would be equally consistent with a converse pattern. If it had turned out, contrary to fact that 6-month-olds succeeded in manual searches while failing to show that they can represent briefly unperceived objects in violation-of-expectation tasks when the mode of disappearance is occlusion, the Core Knowledge View would be no less appealing.

Spelke (1994, 441–2) writes that core knowledge has limited effects on behaviour, being usually manifest in the control of attention (as measured by dishabituation, gaze and looking times) and rarely or never manifest in purposive actions. Of course, this is exactly what needs to be true if invoking core knowledge is to solve the Linking Problem. Or almost exactly—core knowledge would need to be manifest in purposive actions and not gaze when objects disappear by endarkening (see Section 4.3). And maybe this is true. But none of the ways core knowledge is standardly introduced provide any reason to suppose that this is true. And this means that we cannot know whether it is really core knowledge that solves the Linking Problem.

Carey (2009) suggests a further feature of core knowledge: it is an iconic representation, unlike ordinary knowledge. That is, core knowledge has something in common with paper photographs and maps: cutting a photograph down the middle usually gives you two photographs, each of which is a photograph of a smaller part of what the whole photograph was of. In saying that core knowledge is iconic, Carey means that it is structurally similar: a piece of core knowledge has proper parts which are themselves bits of core knowledge about something smaller. This is an interesting idea, but is it relevant to solving the Linking Problem? The problem, once again, is that iconicity doesn't bear on the puzzling pattern of findings. An adequate solution to the Linking Problem should distinguish between the actual world in which 6-month-olds cannot manually search for objects and an alternative, merely possible world in which those infants cannot represent briefly unperceived objects in violation-of-expectation tasks when the mode of disappearance is occlusion. Invoking iconicity would be no less compelling in the alternative, merely possible world. So even knowing whether infants' earliest representations of objects are iconic would not be enough to solve the Linking Problem.

The Core Knowledge View offers an improvement over the idea that 4-month-olds know facts about briefly unperceived objects insofar it does not generate incorrect predictions. But this seems to be mainly because the Core Knowledge View fails to generate any relevant predictions at all.

You might object that two questions are being conflated. One question is the Linking Problem: What links the Principles of Object Perception to infants' minds? On the Core Knowledge View, the answer is that the principles are encoded in core knowledge. Another question is, What explains the puzzling patterns in infants' earliest abilities concerning physical objects? You might object that this is a separate question and not one that a proponent of the Core Knowledge View needs to answer.

It is true that we should distinguish these questions, but they are closely related. How can we tell whether a proposed solution to the Linking Problem is correct? The Linking Problem arises because we want a theory of infants' abilities which is not merely descriptively adequate but also explanatorily adequate (see Section 3.1). The point of solving the Linking Problem is to explain and predict facts about infants' abilities. Success in doing this is what allows us to determine the likely correctness of a proposed solution. As the Core Knowledge View neither explains puzzling patterns in infants' abilities nor generates fine-grained novel predictions, there is no way of knowing whether it solves the Linking Problem.

5.3 How not to define something

The failure of the Core Knowledge View to enable us to solve the Linking Problem is related to the way that core knowledge and the core system are defined. The standard approach to saying what core knowledge and core systems are is essentially just to give a list of features (as we saw in Section 5.1). Can this approach yield a notion of core knowledge that will be useful in explaining cognitive development?

When confronted with a list of features, it is natural to ask why those particular features should 'come as a package' (as Adolphs (2010, 759) and Keren and Schul (2009, 537) both stress in a different context). Why, for instance, should we accept that core systems are both largely innate and largely unchanging through development? What unifies these features? Baillargeon (2002) agrees that infants' abilities to represent objects may be largely innate, but she also argues that they undergo substantial change as a result of learning over the first year of life. There is nothing obviously incoherent about this combination of claims. It is also conceivable that future discoveries could support the opposite combination of views: infants' abilities to represent objects are not innate but a consequence of learning in the first months of life; and, once in place, these are largely unchanging throughout development. As far as we know, there is no compelling reason to accept that the features used to define core systems should all come as a package.

The approach to defining core systems that we have followed means that to claim that infants have core systems is to make a bold conjecture. In general, bolder conjectures are better. But in this case, boldness is merely recklessness. The conjecture is bold merely because multiple claims have been artificially packaged together. It is not a theory about the architecture of the mind: it is a parlay bet.

I don't mean to imply that there is anything intrinsically wrong with parlay bets. (A parlay bet is one which combines bets on several outcomes and pays out only if all the component bets are right.) After all, such bets can pay off nicely if things go your way. But a parlay bet is not a theory. In invoking core systems and core knowledge, we are aiming to understand something about 'the architecture of the mind' (Carey 2009, 67). So core knowledge is supposed to be an explanatory notion. But no explanatory notion can be adequately characterized merely by listing features. We will not get far in understanding the mind's architecture merely by listing features.

In order to be in a position to invoke core knowledge in solving the Linking Problem and explaining how humans first come to know facts about particular physical objects (and facts in other domains too), we will need a different approach to understanding what core knowledge is.

5.4 Will invoking modularity help?

Our aim is to solve the Linking Problem, that is, to understand what links the principles describing infants' earliest abilities concerning physical objects to their minds. The notion of core knowledge appears to provide a way forward (see Section 5.1). But the standard approach to defining core knowledge (and core systems) makes it unsuitable for constructing explanations (see Section 5.2). This looks like the sort of problem a philosopher should be able to help with.

While philosophers have rarely thought about core knowledge in depth, there has been much discussion of a related notion, namely, modularity. Fodor (1983, 101) characterizes modules as systems which are among other things innate and encapsulated. Both of these properties have been invoked in characterizing core systems too (see Section 5.1). There are further points of similarity between modules and core systems. Fodor says modules are domain-specific. That is, one module might concern physical objects while another would concern agency and action, say. This is exactly how core systems are thought of.

In some ways, Fodor's notion of module is broader than that of core system. According to Fodor, modules are 'systems whose operations present the world to thought' (1983, 40). They therefore include perceptual systems, which are not usually regarded as core systems. But as Spelke and Carey introduce core systems by comparison with perceptual systems (Carey and Spelke 1996, 517; Spelke 2003, 278), this difference is plausibly terminological.

Despite the overlap in features used to characterize core systems and modules, there are also some differences. Fodor stipulates that modules are computational systems and that they exhibit neural specificity (so can be identified neurophysiologically). For their part, Spelke and Carey specify that core systems arise from systems already present in the evolutionary ancestors of modern humans (see Section 5.1).

The pattern of similarities and differences makes it hard to know how modules relate to core systems. Some appear to hold, roughly, that core systems are modules (Spelke et al. 1995, 137), or that the account of core systems is a revision of Fodor's account of modularity (Hermer and Spelke 1996). It has also been suggested that core systems compete with modules in the sense that the mind can contain one or the other but not both (Spelke 1988); and, alternatively, that core systems might exist alongside modules (Carey 1995). As things stand, my sense is that there is not so much evidence concerning the features associated with core systems and modules that we can make empirically motivated distinctions between multiple different theories in this area. (To illustrate, see Section 9.3 on evidence for innateness.) I therefore provisionally take 'module' and 'core system' to be different labels for a single thing.

Overall, existing theoretical discussions of modularity do not seem help with the Linking Problem. Modules are standardly defined the same way that core systems are standardly defined: by stipulating a collection of features, none of which allow us to explain the puzzling patterns in infants' abilities concerning physical objects (see Table 4.1). For our purpose, these accounts of modularity share the defects of the standard account of core systems. To recap (see Section 5.2), none of the features associated with core knowledge or modularity explain the actual pattern of discrepancies in 6-month-olds' abilities to track briefly unperceived objects. Exactly the same features could be invoked in a world in which there were a different pattern of discrepancies. This is related to the way the notions of core knowledge and modularity are introduced: by merely listing features, we are making a parlay bet where we should be providing a theory. If we want to solve the Linking Problem, we will need a deeper understanding.

5.5 Conclusion

One leading approach to solving the Linking Problem hinges on the notion of core knowledge (Section 5.1). According to the Core Knowledge View, infants' earliest representations of physical objects are representations in core systems.

This view is initially promising insofar as it does not generate the same incorrect predictions that the Simple View does. But it also fails to explain what we need to explain, and to generate the predictions we need to generate, if we are to know we have solved the Linking Problem. As we saw in Chapter 4, there are surprising patterns when infants of different ages manifest abilities to represent objects as persisting (see Sections 4.1 and 4.2). Appealing to a general theory core knowledge or modularity will not enable us to explain why infants manifest abilities to represent objects as persisting when reaching in the dark but not when reaching behind a barrier. Existing theories of core knowledge are equally compatible with the converse pattern. Nor do these theories generate fine-grained, readily testable new predictions about the conditions under which infants will manifest their abilities to represent objects as persisting.

What next? There are many other general theories about states and processes that are neither mindless nor like knowledge, belief and other common-sense psychological states (including, for example, Stich (1978) on subdoxastic states and Cohen (1992) on belief versus acceptance). My guess, though, is that these general theories will prove likewise unable to provide the kind of explanations and to generate the sort of predictions we need. Instead of elaborating on a general theory that is supposed to work for any domain of knowledge, we should consider in more detail what is known about object cognition in adults.

Notes

6 Object indexes and motor representations of objects

Four-month-old infants have unexpectedly rich abilities concerning physical objects. They can segment objects, represent them as persisting and track their causal interactions; and their abilities to do all three things can be described by a single set of principles about how physical objects behave, the Principles of Object Perception (as we saw in Chapters 2 and 3). How are these early abilities related to knowledge of physical objects? An initially promising conjecture with radical implications for understanding the origins of knowledge is that the principles are things infants know or believe. But this conjecture systematically generates incorrect predictions (as we saw in Chapter 4). On rejecting it, we are immediately confronted with the Linking Problem. If not knowledge or belief, what does link principles describing infants' abilities concerning physical objects to their minds, and thereby explains their abilities?

The most influential and best developed attempt to answer this question invokes core knowledge (see Chapter 5). But as introduced so far, merely invoking core knowledge does not seem to generate the kind of predictions that would enable us to know whether it is true, nor does it explain apparently contradictory patterns in infants' abilities. There may be nothing wrong with invoking core knowledge, but doing so does not enable us to know we have solved the Linking Problem.

In this chapter we will develop a conjecture which may explain the apparently contradictory patterns of evidence about infants' abilities concerning physical objects, and which might almost enable us to solve the Linking Problem.

As a first step we will consider a brilliant conjecture about the cognitive mechanisms underpinning infants' abilities concerning physical objects. Several researchers have conjectured that infants' abilities concerning physical objects are a consequence of something called object indexes (Leslie et al. 1998; Scholl and Leslie 1999; Carey and Xu 2001; Scholl 2007). In developing a version of their view, this chapter will also introduce the method of signature limits. By the end of this chapter you should have a sense of how these ingredients, object indexes and signature limits, can be used to construct developmental theories with readily testable predictions which provide an alternative to, or refinement of, theories relying on knowledge, belief or core knowledge. But first, what are object indexes and why should we believe they exist?

6.1 Object indexes in adult humans

Start with an analogy. An old-fashioned logistician is keeping track of supply trucks by sticking pins in a map. She updates the positions of the pins regularly, based on her information about the terrain and new information about the trucks' fates. Because she only has partial information, much of this is guesswork. But she does not guess at random. Instead she relies on certain principles. For instance, she operates in ways that presuppose the trucks will move along continuous paths, and that they will tend to keep going in roughly the direction they are currently moving.

The system of pins and principles provides a way of tracking trucks. The pins are not representations in the sense that cargo manifests or photographs of the trucks are. Instead the idea is that behaviours of the pins correspond to the behaviours of the trucks. For each pin there is a corresponding truck and, ideally, the movements of the pin reflect the movements of the truck it corresponds to. (Of course the logistician will sometimes stop tracking a truck and assign its pin to a new truck.)

Object indexes are a mental counterpart of the pins: they are things that point to, or index, objects. Like the pins, object indexes are part of a system for tracking objects' movements. And like the logistician, this system is able to function well enough on the basis of partial information because it operates in ways that presuppose principles about physical objects. For instance, it operates in ways that presuppose objects move along continuous paths.

Why believe that object indexes exist in adult humans? One reason is that they can track at least four moving objects simultaneously. (There is debate about exactly how many objects can be tracked simultaneously. Some say the limit is four while others claim eight; see Alvarez and Franconeri 2007.) Suppose you are shown a display involving eight stationary circles (see Figure 6.1). Four of these circles flash, indicating that you should track these circles. All eight circles now begin to move around rapidly, and keep moving unpredictably for some time. Then they stop and one of the circles flashes. Your task is to say whether the flashing circle is one you were supposed to track. Adults are good at this task (Pylyshyn and Storm 1988), indicating that they can use at least four object indexes simultaneously.

The task just mentioned is called multiple object tracking, or MOT for short. It has been widely studied and there are several theories about the mechanisms which underlie humans' abilities to track multiple, simultaneously moving objects. If you look into this research, you will find people talking about 'FINSTs', 'object files' and the like. Although the details are daunting and the differences between theories subtle, we need consider only points on which most theories agree. Most importantly for us, there is quite wide agreement that multiple object tracking is made possible by a system of object indexes analogous to the logistician's pins.

Figure 6.1 A multiple object tracking task. You first see eight circles (at t = 1), four of which flash (at t = 2). Then all circles move around unpredictably for some time (at t = 3). Finally, one circle flashes (at t = 4). Your task is to say whether this last circle is one of the four which flashed at the start.

Source: Pylyshyn (2001), Figure 6.

Another kind of evidence for the existence of a system of object indexes in adult humans arises from something called the object-specific preview benefit. Suppose that once again you are shown an array of objects (see Figure 6.2). At the start a letter appears briefly on each object. (It is not important that letters are used; in theory, any readily distinguishable features should work.) The objects now start moving. At the end of the task, a letter appears on one of the objects. Your task is to say whether this letter is one of the letters that appeared at the start or whether it is a new letter. Consider just those cases in which the answer is yes: the letter at the end is one of those which you saw at the start. Of interest is how long this takes you to respond in two cases: when the letter appears on the same object at the start and end, and, in contrast, when the letter appears on one object at the start and a different object at the end. It turns out that most people can answer the question more quickly in the first case. That is, they are faster when a letter appears on the same object twice than when it appears on two different objects (Kahneman, Treisman and Gibbs 1992). This difference in response times is the object-specific preview benefit. Its existence shows that, in this task, you are keeping track of which object is which as they move. This is why the existence of an object-specific preview benefit is taken to be evidence that object indexes exist.

Is the system of object indexes that explains adults' abilities to track multiple objects the same system of object indexes that explains the object-specific preview benefit? I think this is an open question.301 Fortunately we do not need to answer it. Having seen what object indexes are and the evidence that one or more systems of object indexes exist in adults, our question is how this might bear on infants' abilities concerning physical objects.

6.2 Object indexes and the principles of object perception

Infants have sophisticated abilities concerning physical objects which are manifest from 4 months of age or earlier (see Chapters 2 and 3). But what kind of mechanism is responsible for these abilities? Could it be object indexes? Could infants' abilities concerning physical objects be wholly or in part a consequence of the fact that a system of object indexes exists in humans from the first few months of life?

Figure 6.2 A demonstration of the object-specific preview benefit

Source: Kahneman, Treisman, and Gibbs (1992), Figure 3.

To answer this question, we need to consider infants' various abilities individually. Infants can segment objects, represent them as persisting and track their causal interactions. Which if any of these are even potentially a consequence of having a system of object indexes?

Take segmentation first. Four-month-olds can segment objects, that is, they can identify where one object ends and another begins. Moreover they make appropriate use of both featural information and objects' movements in segmenting objects (see Sections 2.2 and 2.3).

Consider the possibility that this ability to segment physical objects is a consequence of two things. First, insofar as infants use featural information, they are relying on visual and other perceptual processes that may be only indirectly related to any system of object indexes. Second, infants' use of information about movements to segment physical objects is a consequence of constraints on how object indexes are assigned: things that appear to be moving independently cannot be assigned the same index, whereas things that move together are typically assigned the same index. To illustrate, recall Spelke, Hofsten and Kestenbaum's (1989) experiments depicted in Figure 2.8. In these experiments, infants who see two shapes moving together treat them as parts of a single object, whereas infants who see the two shapes moving in opposite directions treat them as parts of distinct objects. This might in principle be because infants have a system of object indexes which assigns one index to the shapes when they move together and two indexes to the shapes when they move in opposite directions.

So much for segmentation. What about infants' abilities to represent objects as persisting? For instance, consider what happens when infants see an object disappear completely behind a barrier. As mentioned in Section 3.2, 4-month-old infants, like adults, manifest an expectation about the object's reappearance by looking in anticipation to the far side of the barrier. This suggests that infants continue to represent the object (and its velocity) even while it cannot be seen. Could this and other ways in which infants can represent objects as persisting likewise be a consequence of a system of object indexes?

To answer this question, we need to ask one about object indexes. What happens to an object index when the object it indexes disappears behind a screen? To find out, we can use the object-specific preview benefit. Whereas the object-specific preview benefit was initially used in arguing that object indexes exist, we can now reverse direction and use the presence of a particular object-specific preview benefit as a marker of a particular object index. Suppose a shape disappears on one side of a screen as if travelling behind it and then, some time later, a shape appears on the other side of the screen as if emerging from it. We want to know whether the shapes are assigned the same object index. We can find this out by determining whether there is an object-specific preview benefit for the two shapes. If there is, we can conclude that they are assigned the same object index. Using this method it is possible to determine that object indexes can index objects while they are completely hidden by a screen. Since an object can have the same object index before and after disappearing, it is conceivable that some abilities to represent objects as persisting could be a consequence of these object indexes.

Is it plausible that infants' abilities to represent objects as persisting could be entirely a consequence of their having a system of object indexes? This would require not merely that object indexes could in principle allow infants to represent objects as persisting, but also that infants represent an object as persisting exactly when it is indexed by a single object index. Infants' abilities to represent objects as persisting are partially characterized by a set of principles, the Principles of Object Perception (these are listed in Table 3.1). So if infants' abilities to represent objects as persisting are entirely a consequence of their having a system of object indexes, the same principles should characterize the operation of object indexes. Do they?

It turns out that they probably do (Mitroff, Arita and Fleck 2009; Scholl and Flombaum 2010). Let me explain by returning to the old-fashioned logistician who is keeping track of supply trucks. In doing this, she has only quite limited information to go on. She receives sporadic reports that a supply truck has been sighted at one or another location. But these reports do not specify which supply truck is at that location. She must therefore work out which pin to move to the newly reported location. In doing this she might rely on assumptions about the trucks' movements being constrained to trace continuous paths, and about the direction and speed of the trucks typically remaining constant. These assumptions allow her to use the sporadic reports that some truck or other is there in forming views about the routes a particular truck has taken. A system of object indexes faces the same problem when the indexed objects are not continuously perceptible. What assumptions or principles are used to determine whether this object at time t1 and that object at time t2 have the same object index pinned to them?

The answer is not completely straightforward. Suppose object indexes are being used in tracking four or more objects simultaneously and one of these objects—call it the first object—disappears behind a barrier. Later two objects appear from behind the barrier, one on the far side of the barrier (call this the far object) and one close to the point where the object disappeared (call this the near object). If the system of object indexes relies on assumptions about speed and direction of movement, then the first object and the far object should be assigned the same object index. But this is not what typically happens. Instead it is likely that the first object and the near object are assigned the same object index.302 If this were what always happened, then we could not fully explain how infants represent objects as persisting by appeal to object indexes because, at least in some cases, infants do use assumptions about speed and direction in interpolating the locations of briefly unperceived objects. There would be a discrepancy between the Principles of Object Perception which characterize how infants represent objects as persisting and the principles that describe how object indexes work.

But this is not the whole story about object indexes. It turns out that object indexes behave differently when just one object is being tracked and the object-specific preview benefit is used to detect them. In this case, it seems that assumptions about continuity and constancy in speed and direction do play a role in determining whether an object at t1 and an object at t2 are assigned the same object indexes (Flombaum and Scholl 2006; Mitroff and Alvarez 2007). In the terms introduced in the previous paragraph, in this case where just one object is being tracked, the first object and the far object are assigned the same object index. This suggests that the principles which govern object indexes may match the principles which characterize how infants represent objects as persisting.

The question we are considering is whether infants' abilities concerning physical objects could be a consequence of their having a system of object indexes. We have seen that infants' ability to segment objects might be the joint upshot of perceptual processes and object indexes. Here the idea is that perceptual processes can explain much of infants' use of featural information in segmenting objects, and object indexes can explain their use of motion information. We have also seen that object indexes can be used in tracking objects which are briefly unperceived, for example, because they have disappeared behind a barrier.

Because the principles describing how object indexes operate appear to match the Principles of Object Perception, which characterize how infants represent objects as persisting, it seems that the existence in infants of a system of object indexes could explain their abilities to represent objects as persisting.

We have not considered whether infants' abilities to track causal interactions among objects might also be a consequence of their having a system of object indexes. I want to skip this question in favour of another.303 So far we have seen some evidence that a system of object indexes could in principle underlie infants' abilities concerning physical objects. Is there any evidence that it does actually underlie them?

6.3 The CLSTX Conjecture

Reflection on object indexes and the Principles of Object Perception in Section 6.2 has led us to a conjecture:

Five-month-olds' abilities to segment objects, to represent them while briefly unperceived and to track their causal interactions are not grounded on belief or knowledge: instead they are consequences of the operations of a system of object indexes.

I will call this the CLSTX Conjecture because versions of it were arrived at, perhaps independently, by Carey, Leslie, Scholl, Tremoulet and Xu.304

The CLSTX Conjecture is a step towards constructing an alternative to the Simple View. The hope is that the CLSTX Conjecture will not generate the incorrect predictions that follow from the Simple View. Even better, if true it will provide us with a way of solving the Linking Problem (see Section 4.4) and perhaps enable us to identify representations that are not knowledge or belief, and so eventually illuminate how humans first come to know simple facts about particular physical objects.

How does the CLSTX Conjecture promise to solve the Linking Problem? According to this conjecture, what links the Principles of Object Perception to an individual mind is a system of object indexes. The Principles are explanatorily adequate and not merely descriptively adequate because they characterize how the system of object indexes operates. This is important for understanding development. Whereas the Simple View entails that infants' earliest abilities already involve beliefs about, or knowledge of, general principles about physical objects and their interactions, the CLSTX Conjecture allows that these early abilities need not involve belief or knowledge at all. Rather than presupposing knowledge, they may therefore provide a foundation for its emergence in development.

What is the evidence for the CLSTX Conjecture? In Section 6.2, we saw that the Principles of Object Perception, which characterize infants' abilities concerning physical objects, resemble principles describing how object indexes work. This indicates formal adequacy: object indexes could in principle explain infants' abilities. But it is not evidence that they actually do so. After all, the resemblance of the Principles of Object Perception to principles describing how object indexes work could be a consequence of the fact that, given the way physical objects actually behave, the same principles will describe many entirely different systems for tracking them.

Is there more direct evidence for the CLSTX Conjecture? Object indexes are probably present in infants by 6 months of age at the latest (Richardson and Kirkham 2004). It is also likely that in infants, as in adults, object indexes are typically maintained when an object disappears behind a barrier. How do we know? In adults there is a pattern of brain activity which appears to be characteristic of processes involved in maintaining an object index for an object that is briefly hidden from view. Kaufman, Csibra and Johnson (2005) asked whether the same is true of infants. They measured brain activity in 6-month-olds infants as they observed a display typical of an object disappearing behind a barrier. They found the pattern of brain activity characteristic of maintaining an object index. This suggests that in infants, as in adults, object indexes can attach to objects that are briefly unperceived. So it is not just conceivable that infants' abilities to represent objects as persisting could be a consequence of object indexes: we also know that object indexes are at least sometimes present when infants are representing objects as persisting.

The evidence we have so far gets us as far as saying, in effect, that someone capable of committing a murder was in the right place at the right time. Can we go beyond such circumstantial evidence? The key to doing so is to exploit signature limits.

6.4 Signature limits

In general, a signature limit of a system is a pattern of behaviour the system exhibits which is both defective given what the system is for and peculiar to that system. To give a simple example, suppose there is a machine for performing addition which generally works well but gives incorrect answers when asked to add twin primes. This is a signature limit of the system. By contrast, the system's giving incorrect answers in very hot conditions, or when very large numbers are input, is not a signature limit. After all, lots of systems fail in these conditions.

Signature limits are useful for determining whether apparently disparate effects are the effects of a single type of system. To continue the simple example, suppose you want to know whether two people are using the same type of machine to do addition. If you ask them both what 2 + 2 is and they both give you the same correct answer, this tells you almost nothing because lots of different machines can reliably give you correct answers. But if you ask them both what the sum of some twin primes is (137 + 139, say) and they both give you the same incorrect answer, you have evidence that they are using the same type of machine.

Perhaps we can use signature limits to identify evidence for the CLSTX Conjecture. Are any signature limits of a system of object indexes also signature limits of whatever underpins infants' abilities concerning physical objects?

Carey (2009) argues that the signature limits of object indexes in adults are reflected in corresponding limits on infants' abilities to track briefly occluded objects. What is her argument?

One signature limit of a system of object indexes is that featural information sometimes fails to influence how objects are assigned in ways that seem quite dramatic. Consider Figure 6.3. A patterned square disappears behind the barrier; later a plain black ring emerges. If you consider speed and direction only, these movements are consistent with there being just one object. But given the distinct shapes and textures of these things, it seems all but certain that there must be two objects. Yet in many cases these two objects will be assigned the same object index (Flombaum and Scholl 2006; Mitroff and Alvarez 2007). So one signature limit of systems of object indexes is that information about speed and distance can override information about shape and texture.

Figure 6.3 A patterned square disappears behind the barrier; later a plain black ring emerges. How likely is it that they are the same object?

Source: Scholl (2007), Figure 4.

Is this also a signature limit of infants' abilities to represent objects as persisting? Xu and Carey (1996) showed 10-month-olds the sequence in which a duck and a ball appear in sequence. Infants first see a yellow rubber duck appear from behind a screen and then disappear behind it again. They then see a white foam ball appear from behind the screen and return behind it. If infants can use featural information like colour, shape and texture in representing objects as persisting, then they should expect there to be two objects behind the screen, just as adults do. In fact, even 10-month-olds (who have probably been representing objects as persisting for at least six months) appear to have no such expectation. This is not a function of the particular objects used; changing them to entirely different objects makes no difference. Xu and Carey offer an impressively deadpan report on a failure that is astonishing from the point of view of an adult:

babies failed to demonstrate that they could use the differences between a yellow rubber toy duck emerging from one side of the screen and a white Styrofoam ball emerging from the other side of the screen to infer that there must be at least two objects behind the screen. (1996, 129)

This failure is striking because infants do generally show an ability to make inferences like this in the paradigm that Xu and Carey used. Apparently infants' early developing abilities to represent objects as persisting have a signature limit matching a signature limit of a system of object indexes, and that this signature limit is not easily overcome (Carey and Xu 2001; Scholl 2007; Carey 2009, 77–83).

How strong is this argument? As it hinges on evidence from a single study, we should be cautious in putting too much weight on it. More interestingly, the argument is complicated by evidence that infants around 10 months of age do not always fail to use featural information appropriately in representing objects as persisting (Wilcox and Chapa 2002). In fact, McCurry, Wilcox and Woods (2009) report evidence that even 5-month-olds can make use of featural information in representing objects as persisting (see also Wilcox 1999). Likewise, object indexes are not always updated in ways that amount to ignoring featural information (Hollingworth and Franconeri 2009; Moore, Stephens and Hein 2010). It remains to be seen whether there is really an exact match between the signature limit on object indexes and the signature limit on 4-month-olds' abilities to represent objects as persisting. The CLSTX Conjecture is a bet on the match being exact.

Other signature limits characterize the operation of object indexes. For instance, there are limits on the number of objects that can be tracked simultaneously (perhaps because there are a limited number of pins or object indexes); it is possible that there is a matching limits in infants' abilities to represent objects as persisting (Cheries, Wynn and Scholl 2006). There are also limits on how long object indexes can be maintained. For instance, object indexes associated with the object-specific preview benefit can last for up to 8 seconds in adult humans (Noles, Scholl and Mitroff 2005). It is conceivable that 4-month-olds' abilities to represent objects as persisting are also subject to a similar limit.305 We also know that infants' initial abilities to represent objects as persisting can be affected by having objects follow untypical trajectories (Bremner, Slater and Johnson 2015, 5) and by the way objects disappear from view (Charles and Rivera 2009, 1004). It is possible that there are signature limits here, and conceivable that these correspond to signature limits of a system of object indexes. As more and more evidence concerning signature limits appears, we can be increasingly confident that infants' abilities concerning physical objects are at least in part a consequence of a system of object indexes, as the CLSTX Conjecture claims.

6.5 Knowledge or core knowledge or ...?

The CLSTX Conjecture has an advantage over the Simple View which is not widely recognized. This is that object indexes are independent of beliefs and knowledge states. Having an object index pointing to a location is not the same thing as believing that an object is there. And nor is having an object index pointing to a series of locations over time the same thing as believing or knowing that these locations are points on the path of a single object. Further, the assignments of object indexes do not invariably give rise to beliefs; nor need such assignments match your beliefs (Mitroff, Scholl and Wynn 2005; Scholl 2007). To emphasize this point, consider once more this scenario in which a patterned square disappears behind the barrier; later a plain black ring emerges (as depicted in Figure 6.3). You probably don't believe that they are the same object, but they probably do get assigned the same object index. Your beliefs and assignments of object indexes are inconsistent in this sense: the world cannot be such that both are correct.

The independence of object indexes from beliefs means that the CLSTX Conjecture does not entail that 4- and 5-month-olds have beliefs or knowledge about physical objects. It does not generate the Simple View's incorrect predictions about 4- and 5-month-olds' abilities concerning physical objects. In between mindless behaviour and propositional attitudes like belief and knowledge, there are things such as object indexes.

Compare the CLSTX Conjecture with the Core Knowledge View. We saw that mere appeal to core knowledge fails to generate useful predictions in Chapter 5. By contrast, invoking object indexes and their signature limits generates many readily testable predictions. But should we think of the CLSTX Conjecture as a competitor to the Core Knowledge View? Or is it better understood as a development of this View?

Carey (2009, Chapter 3) presents the CLSTX Conjecture as a development of a view about Core Knowledge. This makes sense insofar as a system of object indexes may have some of the features used in characterizing core knowledge, such as information encapsulation. And a system of object indexes is probably largely unchanging through development, just as core knowledge is supposed to be.

There is a cost to identifying object indexes with core knowledge, however. According to Carey, core knowledge is a 'type of conceptual structure ... that differs systematically from ... sensory/perceptual representation' (2009, 10; see Section 5.1, for discussion). Since object indexes are a broadly perceptual phenomenon, this claim appears to be in tension with thinking of the CLSTX Conjecture as a development of the Core Knowledge View.

Whether or not we should regard a system of object indexes as a core system (see Section 8.2, for more on this), we should not neglect a key virtue of the CLSTX Conjecture: like the Simple View, it does not require us to postulate novel kinds of representation or knowledge. Instead it is an attempt to explain infant cognition and its development by appealing to systems that are already required for understanding adults' abilities. As we will see in every domain of knowledge, this strategy appears to be the winning one. Much of the progress researchers have made in understanding the emergence of knowledge in development comes from identifying signature limits on infants' capacities and making connections between these capacities and cognitive systems that are already relatively well understood in human adults.

The discussion so far assumes that the CLSTX Conjecture is correct. Just here we face an important complication: the Conjecture is false.

6.6 Against the CLSTX Conjecture

We have seen that there is some evidence for the CLSTX Conjecture, according to which infants' abilities concerning physical objects are a consequence of their having a system of object indexes (see Section 6.4). But can the CLSTX Conjecture explain the puzzling pattern of discrepancies in 4- and 5-month-olds' abilities concerning physical objects?

This puzzling pattern was identified in Chapter 4 and was summarized in Table 4.1. Consider first occlusion events and the difference between violation-of-expectation and manual search measures. Why do 5-month-olds fail to manifest their ability to track briefly occluded objects by initiating searches for them after they have been fully occluded? This is consistent with the CLSTX Conjecture because object indexes are independent of beliefs and do not by themselves support the initiation of action. But what about the difference between occlusion and endarkening on violation-of-expectation experiments? Why do infants show evidence of tracking objects that disappear by occlusion but not by endarkening? Whereas object indexes can be maintained through occlusion events, it may be that endarkening a whole scene interferes with the maintenance of object indexes, perhaps because it involves destroying not only the object but the whole frame of reference. This is merely a speculation (as far as I know). But unless something like this is true, the CLSTX Conjecture will fail to explain the difference between occlusion and endarkening on violation-of-expectation experiments. And even assuming for the sake of argument that endarkening impairs the maintenance of object indexes, there is a more pressing challenge to the CLSTX Conjecture.

Why do infants succeed in searching for momentarily endarkened objects? This finding seems to run directly against the CLSTX Conjecture. Object indexes do not survive endarkening (or so we are assuming); and even if they did, they don't enable you to initiate purposive actions. So the CLSTX Conjecture provides two independent reasons to predict that 5-month-olds will not search for endarkened objects. And yet they do.

In short, the CLSTX Conjecture generates predictions we already know to be incorrect. This fact requires us to abandon or revise the CLSTX Conjecture. Considering how objects can be represented motorically will lead us to a novel way of revising it.

6.7 Motor representations of objects

For adults, representing objects is not always a matter of knowledge, belief or object indexes. They can also represent objects motorically. (See Section 1.3 on the notion of motor representation.)

How an object is represented motorically depends on its affordances. In general, you can only represent an object motorically if you can interact with it (Cardellicchio, Sinigaglia and Costantini 2011). Putting an impenetrable barrier between you and an object—even a translucent barrier—means that you can't interact with it, and so the object is unlikely to be represented motorically (Costantini et al. 2010).

Object indexes and motor representations of objects plausibly have complementary features with respect to different modes of occlusion. Whereas object indexes can survive occlusion but (we are assuming) not endarkening, motor representations survive endarkening but not occlusion events in which an occluder is an impenetrable barrier.306 This is illustrated in Table 6.1, which resembles Table 4.1.

Table 6.1 Complementary features of object indexes and motor representations

Survive occlusionSurvive endarkening
Object indexes
Motor representations✗ (barrier)

Object indexes and motor representations of objects also plausibly differ in the kinds of response they enable. As broadly perceptual phenomena, object indexes function to support the allocation of attention and so plausibly guide looking behaviours, including anticipatory looking; but they probably do not enable the initiation of purposive actions. By contrast, motor representations of objects function to support purposive actions; but there is no reason to suppose that they would cause an infant (or adult) to look in anticipation of an object's reappearance (unless it were the target of an action, perhaps). So where only object indexes underpin an infant's (or an adult's) responses, we would expect to observe anticipatory looking and perhaps success on violation-of-expectation tasks (but see Section 7.1); but we would not expect to observe manual searching. Conversely, where only motor representations underpin an infant's responses, we would expect to observe manual searching but not necessarily anticipatory looking or success on violation-of-expectation tasks (see Table 6.2).

The contrasts between motor representations and object indexes (as summarized in Table 6.2), and the relation between these and the pattern in infants' abilities concerning physical objects, invite us to consider the possibility that 4- and 5-month-old infants' capacities to track briefly unperceived objects may involve not only object indexes but also motor representations of objects.

6.8 Conjecture O

Given that objects can be represented motorically, what might we guess about infants' earliest abilities concerning physical objects? Consider Conjecture O, a revision of the CLSTX Conjecture:

Five-month-olds' abilities to segment objects, to represent them while briefly unperceived and to track their causal interactions are not grounded on belief or knowledge: instead they are consequences of the operations of a system of object indexes ...

... and of a further, independent capacity to track physical objects which involves motor representations and processes.307

Conjecture O generates some readily testable predictions. It predicts that impairing or boosting 4- and 5-month-old infants' abilities to represent objects motorically will modulate performance when they are tasked with initiating searches for briefly unperceived objects. Such manipulations should, however, have no effect on these infants' performance on standard anticipatory looking or violation-of-expectation tasks.

Table 6.2 Object indexes and motor representations support complementary types of response

Enable success in violation-of-expectation tasksEnable initiation of purposive action
Object indexes
Motor representations

How could this prediction be tested? Occlusion events often involve screens which are also barriers to action, so that occluding an object is confounded with removing it from the space of action. We need barriers which do not occlude and occluders which are not barriers to action. Suppose we showed 5-month-olds an object disappearing behind an occluder which was no barrier to action. Conjecture O predicts that these infants should search for the object in this case (and, further, that their search behaviours should not be subject to signature limits associated with object indexes).

If a philosopher makes a prediction, she will invariably be able find a psychological study from the past which appears to support it. I am no exception.

McCurry, Wilcox and Woods (2009) tested 5-month-olds using a fringed screen through which you can reach but not see. This is the perfect way to separate barriers to action from occluders. In a familiarization phase, infants were encouraged to reach through the screen. If they did not, their hands were forcibly moved through the screen. In the test phase, they contrasted two events. In one, a red cube was moved behind the screen and, shortly afterwards, a red cube was moved out from the other side of the screen. This was done in such a way as to suggest to adults that a single red cube had moved behind the screen and back out again. In the other event, a red cube was moved behind the screen and, shortly afterwards, a green spotty was moved out from behind the screen. This should indicate to naive adults that the red cube is still behind the screen. Following each event, McCurry, Wilcox and Woods (2009) gave infants an opportunity to search behind the screen. They were interested in whether infants would reach more frequently towards the screen after the red-cube/green-ball event. And so they did.

Why did infants do this? On the CLSTX Conjecture, this pattern of behaviour makes no sense. Object indexes are not sufficient to initiate purposive actions. And, as we saw, a signature limit of object indexes is their disregard for featural information (Carey and Xu 2001). So the CLSTX Conjecture provides two reasons for predicting, incorrectly, that 5-month-olds should not search longer after the red-cube/green-ball event. But understood in terms of motor representations, infants' performance makes perfect sense. The screen is an occluder that is no barrier for action, so no obstacle to representing the objects motorically. And of course motor processes do not typically disregard featural information: they care deeply about the shapes of things, so we should expect a difference between the ball and the cube. McCurry, Wilcox and Woods (2009) have designed the perfect experiment to test Conjecture O.

We should not take McCurry, Wilcox and Woods' findings as providing evidence for Conjecture O, of course. My interpretation of them is entirely post hoc. McCurry, Wilcox and Woods do not interpret them in this way.308 But their experiment does illustrate how predictions of Conjecture O could be tested.

Conjecture O generates many further predictions, meaning it should be readily refutable if false. It predicts that, in the first six months of life, infants' abilities concerning physical objects should show signature limits of object indexes and signature limits of motor representations. Where objects cannot be represented motorically (because they are behind an impenetrable barrier, for example), infants' abilities should be insensitive to featural information to the extent that object indexes are. And where objects cannot be tracked using object indexes (because their disappearance is due to endarkening, for example), infants' abilities should be sensitive to whether the objects are reachable to the extent that motor representations are.

Can Conjecture O explain the puzzling pattern of discrepancies in 4- and 5-month-olds' abilities to track briefly unperceived objects (see Table 4.1)? The gist of how it might do so is suggested by comparing these discrepancies with the complementary features of object indexes and motor representations summarized in Table 6.1. Object indexes should survive objects' disappearing behind occluders (including barriers) but not events like endarkening which disrupt perceptual frames of reference. By contrast, motor representations should survive objects' disappearing due to endarkening (you can reach for, and grasp, objects in the dark) but not their displacement behind barriers.

While Conjecture O is yet to be tested, it does generate some readily testable predictions and has the potential to explain the puzzling pattern of discrepancies in 4- and 5-month-olds' abilities to track briefly unperceived objects.

6.9 Conclusion: paradox lost

Recall the problem we face. Many studies provide evidence that, from around 4 months of age, infants can represent objects as persisting and track their causal interactions (see Chapters 2 and 3). These studies involve a variety of methods, including habituation, violation-of-expectation, anticipatory looking and manual search (as we saw in Sections 2.2, 2.3, 3.2 and 3.4). The studies suggest that infants' abilities can be described by a set of principles setting out how objects behave, the Principles of Object Perception. But what links these principles to infants' minds?

An initially tempting position is the Simple View: the principles specify things which infants know or believe, and they acquire beliefs about particular objects by making inferences from these principles. Yet there is compelling evidence, also from studies involving a variety of methods, that infants do not know about unperceived objects' locations or about causal interactions among objects until months or years later (as we saw in Sections 4.1–4.3). This evidence arises from cases in which groups of infants fail to do things which anyone with their abilities who had the relevant knowledge could hardly fail to do. It provides strong reasons to reject the Simple View. The problem we face is to find an alternative (see Chapter 4). This is the Linking Problem. If not belief or knowledge, what does link the principles describing infants' abilities to their minds?

Conjecture O suggests a solution to the Linking Problem. Four- and 5-month-old infants do not have knowledge of, or beliefs about, particular physical objects. Instead they have two things—motor representations and a system of object indexes—which enable them to segment objects, represent them as persisting and track their causal interactions. So whereas it might look as if an infant knows or believes that a particular object is behind an impenetrable screen, the truth may be that a system of object indexes in the infant has assigned an index to that object and the index points to a location behind the screen. Or, if the screen is no barrier to action, it may be that she represents the object behind it motorically. What connects the Principles of Object Perception to an individual mind is just this: a system of object indexes in the individual, and her motor system, operate broadly in conformity with the principles.

Our new view, Conjecture O, is quite different from the Simple View. Whereas on the Simple View infants have beliefs about, or knowledge of, the Principles of Object Perception, the new view does not entail that any general principles are represented at all. And whereas, on the Simple View, 4-month-old infants believe or know simple facts about particular physical objects, the new view attributes infants no such states. This is why the new view does not generate the incorrect predictions which all but force us to reject the Simple View.

If untrue, Conjecture O should not be difficult to refute because it generates many testable predictions (as we saw in Section 6.8). It predicts that signature limits of object indexes, and of motor representation, will be manifest in infants' abilities concerning physical objects in roughly the first six months of life.

The pattern of findings that gives rise to the Linking Problem is sometimes called a paradox (for example, Meltzoff and Moore 1998, 202). Strictly speaking, there is no paradox here. To get a paradox we at least need to add the further assumption that only knowledge or belief could underpin infants' abilities concerning physical objects. This assumption is related to one about adults: all mental phenomena and their intentional effects can be explained by appeal to belief, knowledge and other propositional attitudes. You can see this assumption at work in Davidson's philosophical inquiries. He attempts to characterize mental phenomena only by appeal to propositional attitudes like belief, desire and intention. Thus pride is a propositional attitude (Davidson 1976), and perception is either a form of belief or a merely causal process (Davidson 1999a, 730). But the discovery of object indexes shows that recognizing things in between merely mindless behaviour and propositional attitudes is essential not only for understanding infants' minds but also for understanding adults' minds too. We do not 'lack ... a satisfactory vocabulary' for describing phenomena in between mindless nature and propositional attitudes (contra Davidson 2001, 127–8, discussed in Section 4.4). Far from it. Discoveries about object indexes, and about motor representation, have led not only to the vocabulary but even to theoretically coherent and empirically testable hypotheses. Belief, knowledge and other propositional attitudes are merely a fraction of the mental phenomena. Only ignoring this makes it tempting to think there might be a paradox.

Accepting Conjecture O does not entail rejecting the Core Knowledge View. The characterization of core knowledge (see Section 5.1) is open enough to allow that having core knowledge of objects might consist in having two things, namely, a system of object indexes and a capacity to represent objects motorically. As we saw in Section 6.5, this means abandoning Carey's idea that core knowledge is a novel type of conceptual structure. Instead, postulating core knowledge does not (for all we have yet seen) require going beyond the crude picture on which the mind comprises epistemic, perceptual and motoric states (see Section 1.3). More importantly, accepting Conjecture O without rejecting the Core Knowledge View entails recognizing that core knowledge of objects lacks unity: it is a hybrid phenomenon, comprising at least two distinct systems. It also entails recognizing that core knowledge generally is likely to lack uniformity across domains: whereas core knowledge of physical objects consists in object indexes plus motor representations of objects, core knowledge in other domains will surely involve quite different kinds of representations and processes.

What does Conjecture O imply about how humans first come to know simple facts about particular physical objects? On the Simple View, such knowledge is already in place by four months of age. By contrast, if Conjecture O is right, then it may be months or years later that knowledge about particular physical objects first appears in development. This makes it coherent to guess that there are multiple foundations for such knowledge. One is having a system of object indexes, another is having a capacity to represent objects motorically. Getting from these to knowledge of physical objects may involve social interaction about objects, perhaps including learning to use tools.

Can we discover more about the transition from infants' earliest abilities concerning physical objects to the acquisition of knowledge of simple facts about physical objects? Doing so depends on identifying a role for metacognitive feelings.

Notes

7 Metacognitive feelings

How do humans first come to know simple facts about particular physical objects? It seems we are still quite far from an answer to this question, the one from the very start of Chapter 2. Evidence incompatible with the Simple View suggests that infants probably do not already know any simple facts about physical objects at 4–6 months of age (see Chapter 4). And considerations in favour of Conjecture O suggest that motor representations of objects and a system of object indexes may both somehow play a role in explaining the developmental emergence of knowledge (see Chapter 6). But how could they do this?

In this chapter, we will approach this question in characteristically destructive philosophical mode: the first step is to see why Conjecture O does not quite explain what it is supposed to explain. We will then explore one way to fill the resulting explanatory gap by invoking something called metacognitive feelings. By the end you should understand what metacognitive feelings are, how they might be linked to the operations of object indexes and why metacognitive feelings might be essential for solving the Linking Problem.

7.1 Objection to Conjecture O

We have been working on the assumption that either the CLSTX Conjecture or, more likely, Conjecture O can explain 4-month-olds' performance on violation-of-expectation tasks involving occlusion. According to either conjecture, their performance on these tasks is a consequence of the operations of object indexes. But what can object indexes explain? The primary functions of object indexes include influencing the allocation of attention and perhaps guiding ongoing action (see Section 6.1). If this is correct, it may be possible to explain anticipatory looking directly by appeal to the operations of object indexes. But operations involving object indexes cannot directly explain differences in how novel things are to an infant. And nor can operations involving object indexes directly explain why infants look longer at a physically possible event than at a physically impossible event.

In short, it is unclear how the operations involving object indexes could explain the looking behaviours they are supposed to explain according to both the CLSTX Conjecture and Conjecture O.

To illustrate, recall Kellman and Spelke's (1983) experiments depicted in Figures 2.4 and 2.5 in Chapter 2. In part of the experiment, infants are habituated to two stick ends, partially concealed by a screen, moving together. When the habituation phase is over, the screen is removed and infants are either shown two unconnected short sticks or they are shown a single connected stick. Infants look longer at the former. In cases such as this, infants (and adults) assign a single object index to the stick ends moving together (see Section 6.2). Consequently, when the screen is removed and two unconnected sticks are shown, two object indexes will be assigned where before a single object index was assigned. This is true (or at least it must be if the Conjecture O is right). But it cannot explain why infants look longer all by itself. How does a difference in operations involving object indexes result in a difference in looking times?

The difficulty this question poses becomes even greater if we consider the timings involved. Mean looking times in the first test trials are around 10 seconds in one condition versus around 40 seconds in the other (see Figure 2.6). Given that, even in adults, object indexes do not survive occlusion for more than around 8 seconds (compare Noles, Scholl and Mitroff 2005), how could they explain such large differences in extended looking times?

The same question arises for looking times in violation-of-expectation experiments, which also involve measuring differences in looking times. It is a mystery how differences in operations involving object indexes could explain such differences in looking time.

Can we disregard these problems as merely the kind of outstanding detail likely to be associated with any bold conjecture? No. The only known way of supporting Conjecture O (and the CLSTX Conjecture) involves relying on the method of signature limits. This cuts both ways. If it were correct to regard Conjecture O as supported by limits on the nature and timing of infants' responses (as argued in Section 6.4), then it must also be right to reject this conjecture when infants' responses do not appear to conform to signature limits.

At this point it may be tempting to suppose that the operations of object indexes give rise to beliefs or to knowledge states, and that these in turn explain the looking behaviours. We should resist this temptation. Not only because the supposition is false (see Section 6.9). Worse, accepting it would mean we had sacrificed simplicity only to end up generating all of the incorrect predictions associated with the Simple View (see Sections 4.1–4.3). We need an alternative.

None of this means we should give up on Conjecture O yet. Perhaps there is something that connects operations involving object indexes with looking behaviours in habituation and violation-of-expectation experiments. And perhaps finding this connection will give us a clearer understanding of the cognitive systems that comprise core knowledge of objects.

7.2 Metacognitive feelings: a first example

What connects operations involving object indexes to patterns in looking duration? It cannot be beliefs or knowledge states (see Section 7.1). Instead I will suggest that it is something called a metacognitive feeling.

According to Koriat (2000, 150), 'metacognitive feelings ... allow a transition from the implicit-automatic mode to the explicit-controlled mode of operation'. We might guess that the operations involving indexes involved an 'implicit-automatic mode' (whatever exactly that means) and that patterns of looking involve an 'explicit-controlled mode'. If that guess is correct, and if Koriat is correct about the role of metacognitive feelings, then they are just what we need to overcome the objection to Conjecture O.

But what are metacognitive feelings? Rather than giving a definition, I think it is helpful to start with particular cases which have been successfully investigated. These successes will give us confidence that there is something worth defining and will limit what a definition should say.

As a first example of a metacognitive feeling, take the sense of agency. As illustrated in Figure 7.1, it's quite well established that there are feelings of agency. These feelings seem to arise from a number of cues, including comparison between outcomes represented motorically and outcomes detected sensorily, and the fluency of an action selection process (that is, the ease or difficulty involved in selecting one from among several possible actions to perform motorically). The latter can be manipulated by, for example, providing helpful or misleading cues to action (Wenke, Fleming and Haggard 2010; Sidarus, Chambon and Haggard 2013; Sidarus, Vuorre and Haggard 2017).

The sense of agency is relevant to us because it serves to link two largely independent processes concerned with evaluating whether you are the agent of an event. One involves detecting the cues just mentioned; the other involves thinking about how likely it is that you are the agent of an event, perhaps in the light of your background knowledge.

Figure 7.1

Figure 7.1 Cues giving rise to a sense of agency

Source: Adapted from Sidarus and Haggard (2016), Figure 5.

Suppose you are a subject in an experiment and the experimenter asks you whether you felt you were in control of an event. You do not need to go with your feelings. You could think about how likely it is that you are the agent of an event. The right answer may well be to say, 'I don't know, this is a psychological experiment so there's a good chance you were tricking me.' But despite all of the possible ways in which reflection on the question might lead to this or another answer, adults systematically give answers which seem to reflect the fluency of action selection. Why?

It seems that action selection fluency modulates the feeling of agency, and that feeling is associated by adults with being the agent of an event. So the feeling plus association can bias adults' answers to questions about agency. (Of course, there may be cases in which adults' answers do not reflect their feelings of agency.)

So what is this feeling (or 'sense') of agency? First, it is phenomenal rather than epistemic. It is an aspect of the phenomenal character of some experience associated with acting. So we can call it a feeling.

Second, the feeling of agency is metacognitive in the sense that its normal causes include processes which monitor action selection and production. So we can call it a metacognitive feeling.401

7.3 More metacognitive feelings

The feeling of agency is far from the only metacognitive feeling. As a second illustration, consider the feeling of familiarity. If you look at the faces depicted in Figure 7.2, you will hopefully have a feeling of familiarity on seeing at least one of these faces. Having a feeling of familiarity need not involve believing that the thing encountered is familiar. Even if you know for sure that you have never encountered the person depicted (and trust me, you have not), the feeling of familiarity will persist. Nor is the feeling merely a perceptual experience. After all, the familiarity of something to you depends on arbitrary past events. You cannot perceptually experience familiarity any more than you can perceptually experience events from your childhood. So the feeling of familiarity is something different from both belief and perceptual experience.

The faces in Figure 7.2 are composites of famous faces. I chose them to illustrate that the feeling of familiarity is not a consequence of how familiar things actually are. Instead it appears to be a consequence of the degree of fluency with which unconscious processes can identify perceived items (Whittlesea 1993; Whittlesea and Williams 1998). Processing familiar faces is generally fluent: morphing familiar faces can create super-fluency.

Feeling of familiarity are not restricted to faces. Learning a grammar can also generate these. Subjects who have implicitly learned an artificial grammar report feelings of familiarity when they encounter novel stimuli that are part of the learnt grammar (Scott and Dienes 2008).

Adults are not compelled to treat feelings of familiarity as actually being about familiarity. It is possible, for example, to exploit the feeling of familiarity in deciding whether a string can be generated by a certain set of rules (for example, Wan, Dienes and Fu 2008). In doing this you are treating the 'feeling of familiarity' as having nothing to do with familiarity as such.

Figure 7.2

Figure 7.2 The feeling of familiarity is not always caused by the fact of familiarity

The feelings of agency and familiarity are not isolated cases. There is also the feeling you have when someone's eyes are boring into your back, the feeling associated with having a name on the tip of your tongue, the feeling of déjà vu (Brown 2003), the feeling of knowing (Koriat 1993), and more besides. These are all candidate metacognitive feelings.

7.4 What is a metacognitive feeling?

It might be useful to compare metacognitive feelings with the feeling of electricity. Contrast two sensory encounters with a wire. In the first you visually experience the wire as having a certain shape. In the second you receive a mild electric shock from the wire without seeing or touching it.402 The first sensory encounter involves perceptually experiencing a property of the wire whereas the second does not. (If anything is perceptually experienced in receiving the mild electric shock, it is probably your own body.) Yet the electric shock involves rich phenomenology, and its particular phenomenal character depends in part on properties of its cause.

Changes in current result in encounters with different phenomenal characters, so that in principle—if contact with live wires were not so often fatal—you could learn to estimate how much electricity is flowing through the wire.

Neither the feeling of familiarity nor the sensations associated with electric current involve standing in any intentional relation to the properties of familiarity or electricity. But they are things that adults can, and often do, interpret as being informative about familiarity and electricity. I suggest that this is true of metacognitive feelings generally:

Metacognitive feelings are, or involve, aspects of the overall phenomenal character of experiences which their subjects take to be informative about things that are only distantly related (if at all) to the things that those experiences intentionally relate the subject to. And there is no further phenomenal feature of metacognitive feelings.

It follows that metacognitive feelings can lead to beliefs only via associations or further beliefs. They are signs which are open to interpretation by their subjects. Just as coming to associate a feeling of electricity with electric current depends on learning (and may never happen), so connecting metacognitive feelings to familiarity, confidence or anything else can only be a consequence of learning.

As the scientist, you can pick out the feeling of familiarity as (very approximately) that metacognitive feeling which is normally caused by the degree to which certain processes are fluent. But as the subject who has that metacognitive feeling, you do not necessarily know what its typical causes are. This is something you have to work out in whatever ways you work out the causes of any other type of event. In this way, metacognitive feelings may contrast with perceptual experiences. Several philosophers hold that perceptual experiences make corresponding beliefs available.403 This is untrue of metacognitive feelings. Instead, metacognitive feelings are not far from being sensations in approximately Reid's sense (1785b, 1785a). Reid's sensations are monadic properties of events, specifically perceptual events, individuated by their normal causes which alter the overall phenomenal character of those experiences in ways not determined by the contents of the experiences (so two perceptual experiences can have the same content while one has a sensational property which the other lacks).404

Much of the work that metacognitive feelings do is surely independent of what you believe. Irrespective of which if any beliefs they lead to, feelings of agency, familiarity, déjà vu and the rest can arouse interest and modulate effort, causing you to slow down and look longer at events that would otherwise seem unremarkable. They may even lead you to explore ways to recreate that feeling. In all of these ways, metacognitive feelings can cause you to focus on events and tasks in contexts that provide opportunities for learning. And they do this regardless of how, if at all, you interpret them. Whatever your research on the feeling of familiarity leads you to conclude about its causes, its occurrence in you will probably still influence how you look at, and attend to, things. Just as the feeling of electricity does.

Note that the partial characterization of metacognitive feelings I am offering is unlikely to be accepted by all researchers. Are metacognitive feelings really like the feeling of electricity, as I am suggesting? Koriat, who is surely the leading researcher on this topic, proposes

‘metacognitive feelings are mediated by the implicit application of nonanalytic heuristics ... [which] operate below full consciousness, relying on a variety of cues ... [and] affect metacognitive judgments by influencing subjective experience itself’ (2000, 158; see also Koriat 2007, 313–15).

Koriat’s deep and carefully developed theory informs the present discussion of metacognitive feelings. His theory is consistent with the characterization of metacognitive feelings I have offered. Indeed, his suggestion that the ‘processes that take off from subjective experience generally have no access to the processes that have produced that experience in the first place’ (Koriat 2007, 314) is in line with my partial characterization of metacognitive feelings. But many other approaches to characterizing metacognitive feelings are incompatible, at least superficially, with thinking of them as like the feeling of electricity.405

Metacognitive feelings have been quite widely neglected in philosophy and developmental psychology. They are a means by which cognitive processes enable perceivers to acquire dispositions to form beliefs about objects’ properties which are reliably true. Metacognitive feelings provide a low-cost but efficient bridge between non-conscious cognitive processes and conscious reasoning (Koriat 2000).

7.5 A metacognitive feeling of surprise?

As a step towards using metacognitive feelings to resolve the objection to Conjecture O (see Section 7.1), we need to postulate a novel metacognitive feeling, one not included on standard lists of metacognitive feelings.

One kind of surprise has features characteristic of a metacognitive feeling. According to Reisenzein (2000, 271): ‘The intensity of felt surprise is not only influenced by the unexpectedness of the surprising event, but also by the degree of the event’s interference with ongoing mental activity.’ On this view, unexpectedness is not, or not only what generates the feeling of surprise. Rather, it is metacognitive monitoring of mental activity: the less fluently an event is processed, the more surprising it feels. This makes sense, given that, within limits, unexpectedness is reliably correlated with the fluency with which an event can be processed, and given that unexpectedness is in general harder to monitor than processing fluency.406

If, as Reisenzein (2000) proposes, there is a phenomenal consequence of mental interference, then it is a metacognitive feeling (see Section 7.4). Let us call it the metacognitive feeling of surprise.

The metacognitive feeling of surprise contrasts with the feelings of agency and familiarity in one interesting respect. Agency and familiarity are stronger when action selection (see Section 7.2) or identification (see Section 7.3) is more fluent. By contrast, the metacognitive feeling of surprise is stronger when the processing of an event is less fluent. But all three do have in common the key characteristic of a metacognitive feeling: they are all feelings that arise from mental processes monitoring the fluency of mental processes.

The existence of a metacognitive feeling of surprise can explain a mundane observation about magic tricks. Even if you have seen the trick before, and even if you know how it is done, so that what you see is exactly what you expect, the trick may nevertheless retain its

magic. How does this ever happen? The retained magic is the metacognitive feeling of surprise, which occurs because the trick interferes with the smooth operation of broadly perceptual processes in you. This is why familiarity, and even insight, do not always prevent you from feeling the magic.

7.6 Conjecture O^m

I introduced metacognitive feelings on the basis that they would enable us to overcome the objection to Conjecture O from Section 7.1. But how are metacognitive feelings relevant to understanding 4-month-old infants’ abilities concerning physical objects?

I conjecture that violation-of-expectation experiments are not far from magic tricks. Operations involving object indexes can give rise to metacognitive feelings of surprise, and these metacognitive feelings can in turn influence looking behaviours in habituation and violation-of-expectation experiments.

To illustrate this conjecture, recall once more the one stick/two sticks experiment depicted in Chapter 2. Imagine it is you, rather than an infant seeing the stimuli. You might have seen them many times before, so that you know just what to expect when that middle box is removed. Still, the event in which the two sticks is revealed is, like a magic trick, liable to feel surprising in some small way. This is characteristic of the metacognitive feeling of surprise: its normal cause is not, or not only, the unexpectedness of the event but rather the extent to which it interferes with ongoing mental activity. But what mental activity does the event of revealing the two sticks interfere with? During the habituation phase, you assign a single object to the stick behind the box. Then, when the habituation phase is over and the box is removed, you are shown two unconnected short sticks. So whereas you assigned a single object index to the stick ends moving together, when the box is removed, two object indexes are needed. There is an error in your system of object indexes, which is an interference with ongoing mental activity. This interference can give rise to a metacognitive feeling of surprise, which could in turn influence your looking behaviour.

We are now in a position to overcome the objection to Conjecture O introduced in Section 7.1. Recall that, according to Conjecture O, 4-month-olds’ abilities to segment objects, to represent them as persisting and to track their causal interactions variously involve object indexes and motor representations. In particular, object indexes (and not motor representations) are held to explain why 4-month-olds succeed in habituation and violation-of-expectation tasks involving objects occluded by an impenetrable screen (see Sections 6.3 and 6.7). The objection to Conjecture O is that object indexes are fundamentally unsuited to explaining success in these tasks, both because they involve voluntary looking behaviours which object indexes are not thought to explain and also because the timings involved are not on the scale on which object indexes are thought to operate (see Section 7.1).

This is a good objection: Conjecture O really cannot explain the looking behaviours that it needs to explain. But the objection can be overcome by switching to what I will call Conjecture O^m (‘m’ for metacognitive). This is just Conjecture O together with the further conjecture that errors in operations on object indexes (and motor representations) can give rise to

metacognitive feelings of surprise. If Conjecture O^m is correct, it is metacognitive feelings of surprise that connect operations on object indexes to looking behaviours.

Note that this is just a guess—and not one that, as far as I know, anyone else would yet endorse. We have moved well beyond the evidence into the realm of theoretical speculation (although perhaps not as far beyond the evidence as proponents of core knowledge, whose theoretical commitments are considerably bolder; see Sections 5.1 and 5.2). But what is the alternative? If you do not accept that metacognitive feelings can be triggered by operations on object indexes, then you face a problem. You will either need to scrap Conjecture O altogether and find an alternative solution to the Linking Problem. Or else you will have to provide an alternative account of how operations on object indexes influence looking times in habituation and violation-of-expectation experiments.407

While much remains uncertain, three things seem clear. First, the existence in 4-month-old infants of a system of object indexes is likely to explain, in part, many of their abilities concerning physical objects (see Section 6.4). Second, operations involving object indexes cannot by themselves explain patterns of looking duration in habituation and violation-of-expectation experiments (see Section 7.1). Third, it might just be that some operations involving object indexes give rise to metacognitive feelings, which in turn influence looking durations.

If this is correct, we have solved the Linking Problem. The problem was to identify what links the Principles of Object Perception to the mind of an individual (see Section 4.4). The solution hinges on Conjecture O^m. In the individual mind there are object indexes and motor representations; and the Principles characterize the individual’s abilities concerning physical objects insofar as these principles constrain operations on object indexes and motor representations. This solution to the Linking Problem has a variety of consequences for understanding the developmental emergence of knowledge, as we will see in the Conclusion to Part I.

7.7 Metacognitive feelings are intentional isolators

We have revised Conjecture O, adding the further conjecture that operations on object indexes and motor representations can give rise to metacognitive feelings of surprise, which in turn can cause looking behaviours such as those observed in habituation and violation-of-expectation experiments. You might object that this combination of conjectures, Conjecture O^m, generates some of the same incorrect predictions that were fatally generated by the Simple View. (These predictions and the evidence disconfirming them were identified in Sections 4.1–4.3.)

To see how the objection might arise, consider another application of Conjecture O^m. What happens when an infant (or adult) observes an object being occluded by an impenetrable barrier? According to Conjecture O^m, an object index attached to the object is maintained at a location behind the barrier. Further, if the barrier is removed and the object does not appear, the system of object indexes will encounter an error condition: an object index is assigned to nothing. This interference in processing the scene can give rise to a metacognitive feeling of

surprise, which could cause the infant to look longer and be more interested in this event than she would be in another, less magical event. So far this is exactly what we want. But the objection is that the metacognitive feeling involves or entails knowledge of, or belief about, the object’s location. In that case, our view would imply that infants know or believe something about the object’s location. So Conjecture O^m appears to generate the incorrect prediction that even 4-month-olds will manually search behind the impenetrable barrier (see Section 4.1).

This objection is based on a mistake about metacognitive feelings, at least as characterized here (see Section 7.4). It is true, of course, that metacognitive feelings enable adults to transition from perceptual or motor processes to beliefs. When encountering an event or viewing a face, for example, metacognitive feelings of agency or of familiarity may enable you to acquire the belief that the event is an action of yours, or that the face is familiar to you. But there is an important detail in how metacognitive feelings enable you (as an adult) to acquire beliefs.

Metacognitive feelings have no intentional objects, or none that are related to any beliefs that you might ordinarily acquire on the basis of them (see Section 7.4). They therefore serve as intentional isolators. That is, they provide a nonintentional link between two intentional states. The processes involved in action selection (or face processing) involve representations of an action (or face), and the beliefs you acquire also intentionally specify an action (or face). But the metacognitive feeling of agency (or of familiarity) that connects these two has no such intentional object. In acquiring the belief, you have to form a view about what caused the feeling.

Compare being stung by a subtle nettle, where the sting can only be felt some time after brushing against plant. Feeling the stinging sensation, you might look around to see what caused it and identify the likely plant behind you. The stinging sensation is like a metacognitive feeling. Neither has an intentional object that is relevant to the belief you eventually acquire. Instead, acquiring the belief involves identifying a likely cause of the sensation or metacognitive feeling.

The objection to Conjecture O^m arises because it is often so effortless for adults to acquire beliefs on the basis of metacognitive feelings that they rarely notice any inference is needed at all. But once we recognize that metacognitive feelings lead to beliefs only via a process of inference which involves identifying a likely cause of the metacognitive feeling, we can see that the objection rests on a mistake. We know 4- and 5-month-olds are not in a position to make such inferences: if they were, they would have beliefs about, or knowledge of, the locations of briefly occluded physical objects; which they do not (see Sections 4.1 and 4.2). So Conjecture O^m does not generate incorrect predictions about infants’ beliefs about, or knowledge of, particular physical objects.

When adults have a metacognitive feeling of surprise, they are likely to have an idea of why they are surprised. Four- and 5-month-olds also have metacognitive feelings of surprise, but they are unlikely to have any idea why they are surprised. This will eventually be critical for understanding how humans first come to acquire knowledge about particular physical objects in development, because it tells us something about how the early developing capacities concerning physical objects are isolated from epistemic capacities.

7.8 Conclusion

Our overall aim is to understand how humans first acquire knowledge of simple facts about particular physical objects in their development. A key discovery is that infants at around 4 months of age already manifest abilities that seem likely to support, somehow, the emergence of such knowledge (see Chapters 2 and 3). The problem we encountered was to understand the kind of cognitive states and processes involved in these abilities. After ruling out belief and knowledge states (in Chapter 4) and finding that merely invoking core knowledge is at best insufficient (in Chapter 5), we hit on Conjecture O, according to which, infants’ abilities are underpinned by a combination of object indexes and motor representations of objects (see Chapter 6). This conjecture is most promising insofar as it does not generate the incorrect projections that other views do, but it does generate many readily testable predictions, some of which have already been tested. Yet in this chapter, we saw that there is an objection to the conjecture (in Section 7.1). The objection can be overcome by invoking metacognitive feelings.

This objection and its resolution matter in two ways. On the one hand, we cannot claim to have characterized the states and processes involved in infants’ earliest abilities concerning physical objects without overcoming it. On the other hand, overcoming the objection indicates that infants’ earliest abilities concerning physical objects involve a further ingredient, namely, metacognitive feelings. This last ingredient will shortly turn out to be essential for understanding something about how humans first acquire knowledge of simple facts about particular physical objects.

Notes

8 Conclusion to Part I

There are two main theories about the nature of infants’ earliest capacities concerning physical objects. According to one, these capacities involve beliefs about, or knowledge of, general principles governing objects’ behaviours. According to the other, understanding these capacities requires postulating a novel kind of representation (or ‘conceptual structure’), something knowledge-like which is not actually knowledge (Carey 2009, 10; see Section 5.1). Both theories are wrong, or so the arguments of Part I has attempted to show.

Instead, 4-month-olds’ abilities are based on a combination of object indexes, motor representations of objects and metacognitive feelings of surprise. None of these three things necessarily involves having any beliefs about, or knowledge of, physical objects. But nor are they novel kinds of representation—each of the three is familiar from comparatively established theories about adult cognition. Infants’ sophistication is real but does not involve knowledge: there is still so much to be explained about the developmental emergence of knowledge of physical objects. On the other hand, the puzzling patterns of performance infants exhibit do not justify postulating novel kinds of representation. The novel feature of infants’ abilities concerning objects is the way object indexes, motor representations and metacognitive feelings work together.

None of this quite answers the question we started with: how do humans make the amazing transition from not knowing any facts about particular physical objects to knowing some facts? Before evaluating our progress with respect to this question, I first want to return to a question that has been outstanding since Section 2.2.

8.1 What is an expectation?

Much developmental research hinges on assumptions about infants’ expectations and things that surprise them, as we saw throughout Chapters 2 and 3. But what are these things? What is an expectation and what is surprise?

Many philosophers take surprise to involve awareness of your own beliefs. For example, Davidson stipulates that to be surprised that there is no coin in my pocket,

[i]t is not enough that I first believe there is a coin in my pocket, and after emptying my pocket I no longer have this belief. Surprise requires that I be aware of a contrast between what I did believe and what I come to believe. Such awareness ... is a belief about a belief.

(2001, 104)

By contrast, developmental scientists use the term ‘surprise’ with no commitment to any such view. For example, Wang, Baillargeon and Brueckner (2004) stipulate that ‘[t]he term surprise is used here simply ... to denote a state of heightened attention or interest caused by an expectation violation’ (Wang, Baillargeon and Brueckner 2004, 168).

On this view, surprise is a subjective consequence of an expectation being violated and need not involve any awareness that the expectation has been violated. The scientists’ characterization of surprise raises more questions than it answers. What is this state caused by an expectation violation? And what is an expectation, in this context?

Given the argument for Conjecture O^m (see Section 7.6), we can answer these questions by appeal to object indexes, motor representations of objects and metacognitive feelings. Operations involving object indexes or motor representations depend on physical objects appearing to behave roughly in accordance with certain principles. For infants to have an expectation about the ways physical objects behave is just for the thing expected to be required if physical objects are to behave in accordance with these principles. And surprise is a metacognitive feeling which occurs when a violation of this expectation causes an error in the system of object indexes.

This way of characterizing expectations and surprise is at odds with ordinary ways of using the words ‘expectation’ and ‘surprise’. In everyday life, an expectation might be something like a belief: to expect rain is to believe it will rain, or will probably rain, for example. By contrast, the technical notion of expectation used in developmental science is best understood as nothing like a belief. In that technical sense, for an event to be expected by you is just for its occurrence to be accordance with principles characterizing perceptual or motor processes in you. Similarly, surprise in everyday life involves there being something you are surprised that. You are surprised that it did not rain, for example. By contrast, the metacognitive feeling of surprise has no comparable intentional object: it is not surprise about anything (see Section 7.7). Just as you can experience feelings of anger without there being anything you are angry about, so having a metacognitive feeling of surprise need not involve being surprised that anything is the case.

This (admittedly speculative and unorthodox) view about expectations and surprise has a radical consequence. Suppose habituation and violation-of-expectation experiments really do measure metacognitive feelings of surprise, as just suggested. Then these methods should not be regarded as revelatory about what subjects believe or know. Why not? Because all the available evidence about metacognitive feelings indicates that they are triggered by monitoring the fluency of processes like facial recognition, action selection, and the like; none of it points to reflective inferential processes operating on beliefs or knowledge states as potential causes of metacognitive feelings. This suggests that habituation and violation-of-expectation results are likely to reflect what is going on at a deeper level than that at which humans are making inferences and acquiring beliefs or knowledge.

8.2 Core knowledge: a lighter account

One consequence of Conjecture O^m concerns core knowledge. If the conjecture is correct, we must either abandon the claim that infants have core knowledge of objects or else recognize that core knowledge of objects lacks unity, being a composite of at least three things that are, to an interesting degree, independent of each other: object indexes, motor representations of objects and metacognitive feelings of surprise. Which path should we take? This may be a largely terminological matter. Carey, one of those responsible for coining the term ‘core knowledge’, takes the CLSTX Conjecture about object indexes to be consistent with her views on core knowledge. A proponent of core knowledge could take the same attitude towards Conjecture O^m.

Recall that core knowledge is characterized by invoking a list of properties including innateness, encapsulation and the rest (see Section 5.1). The view that core knowledge actually has these properties faces three lines of objection. First, there is no compelling theoretical reason to assume that these properties should hang together (as we saw in Section 5.3). Second, there is scarce evidence that the things comprising core knowledge actually have these properties (for consideration of innateness, see Chapter 9). Third, and most pressing, not all of the properties appear to be relevant to explaining infants’ abilities and the developmental emergence of knowledge (see Section 5.2).

We can avoid all three lines of objection by taking a theoretically lighter approach to characterizing core knowledge. Start with a minimally informative definition of core knowledge. In the domain of objects, let core knowledge of physical objects be whatever it is that underpins infants’ abilities to segment objects, represent them as persisting and track their causal actions. And similarly for other domains. Discover the things that actually underpin the abilities—which, in the case of physical objects, are object indexes, motor representations and metacognitive feelings, at least according to Conjecture O^m (see Chapter 6). Call these constituents of core knowledge. Next, check we are justified in distinguishing core knowledge from knowledge proper. To this end, establish whether processes operating on the constituents are distinct from the inferential processes in which knowledge states feature. They should be distinct in this sense: the conditions that influence whether these processes occur or which outputs they generate, and the conditions under which these process result in

certain kinds of behaviours, do not completely overlap. In the case of physical objects, there is abundant evidence for distinctness (see Table 4.1, for a partial summary). Finally, investigate the constituents of core knowledge to discover whether core knowledge exhibits characteristics such as encapsulation and innateness.

On the theoretically lighter approach, whether core knowledge is innate, exhibits informational encapsulation or arises from phylogenetically old systems is a matter for discovery rather than something assumed in advance.

In the case of physical objects, we do already have evidence that core knowledge exhibits an interesting degree of informational encapsulation. This is particularly clear for one of its constituents, object indexes, assignments of which can conflict with a subject’s beliefs (see Section 6.5). The conjectured role of metacognitive feelings in connecting operations on object indexes (and motor representations) to knowledge states also indicates that inferential integration will be limited or absent. As intentional isolators, metacognitive feelings provide a non-inferential link between core knowledge of physical objects and knowledge proper.

Taking the lighter approach means we should be cautious in drawing conclusions about core knowledge generally. There is no guarantee that, assuming it exists at all, core knowledge in other domains will exhibit the same properties as core knowledge of physical objects.

8.3 Development is rediscovery

How do humans get from the abilities concerning physical objects manifested at 4 and 5 months of age to being fully fledged knowers? How do they make the transition to knowing simple facts about physical objects, such as the fact that there are two mice behind that screen?

Leading contemporary answers to this question rely on what I will call the Assumption of Representational Connections: ‘The transition from not knowing any simple facts about physical objects at all to knowing some such facts involves operations on the contents of core knowledge states, which transform them into (components of) the contents of knowledge states.’

This Assumption is required for Spelke’s suggestion that mature understanding of objects, number and mind derives from core knowledge by virtue of core knowledge representations being assembled (Spelke 2000). It is also required for claims by Leslie and others that modules provide conceptual identifications of their inputs (Leslie 1988); and for Karmiloff-Smith’s representational re-description (Karmiloff-Smith 1992). It is even required for Mandler’s claim that ‘the earliest conceptual functioning consists of a redescription of perceptual structure’ (Mandler 1992).

The view currently under consideration, Conjecture O^m, requires us to reject the Assumption of Representational Connections. On the present view, 4- and 5-month-old infants’ abilities to track briefly unperceived objects depend on three things: (1) a system of object indexes; (2) a capacity to represent objects motorically; and (3) metacognitive feelings. So if there is core knowledge of physical objects at all, it is a hybrid of these three (see Section 8.2). Further,

having metacognitive feelings is not a matter of being intentionally related to physical objects (other than perhaps oneself, of course). And only metacognitive feelings and other intentional isolators link operations on object indexes to beliefs and knowledge states. If this is right, it is not true that operations on the contents of core knowledge states somehow transform them into components of knowledge states. Metacognitive feelings insulate any representational features of core knowledge from intentional features of beliefs and knowledge states. The Assumption of Representational Connections must therefore be rejected.

This makes the question about development particularly difficult to answer. It means that rather than being a matter of assembling or re-describing representations, development must be a process of rediscovery. Coming to know simple facts about physical objects is a matter of rediscovering things that are already core knowledge, or at least already implicit in the operations of a system of object indexes.

Some might object that development could not require such rediscovery because it would be hopelessly inefficient to require things already encoded to be learnt anew. But rediscovery is an elegant solution to a practical problem. If you are building a survival system, you want quick and dirty heuristics that are good enough to keep it alive: you don’t necessarily care about the truth. If, by contrast, you are building a thinker, you want her to be able to think things that are true irrespective of their survival value. This cuts two ways. On the one hand, you want the thinker’s thoughts not to be constrained by heuristics that ensure her survival. On the other hand, in allowing the thinker freedom to pursue the truth, there is an excellent chance she will end up profoundly mistaken or deeply confused about the nature of physical objects. So you don’t want thought contaminated by survival heuristics and you don’t want survival heuristics contaminated by thought. Or if some contamination is inevitable, you at least want to limit it. You want beliefs and knowledge states to be inferentially isolated from survival heuristics, including any core knowledge. This is beautifully achieved by giving your thinker systems for tracking objects and their interactions which appear early in development, and also a mind which allows her to acquire knowledge of physical objects gradually over months or years, taking advantage of interactions with objects as well as social interactions about objects—providing, of course, that the two are not directly connected but rather linked only very loosely, via intentional isolators like metacognitive feelings.

8.4 How does rediscovery occur?

Discoveries that information about particular physical objects is already somehow carried in 4-month-olds’ core knowledge make it initially tempting to guess that it will be easy to get from core knowledge to knowledge proper. But matters are complicated by the conjecture that only intentional isolators such as metacognitive feelings link core knowledge to beliefs (at least in the case of physical objects).

The step from core knowledge to knowledge proper is like the step from feeling electric shocks to understanding electricity. Receiving electric shocks might alert you that there is something to be discovered but does not reveal much about the nature of electricity. Like electric shocks, metacognitive feelings of surprise may function to trigger ‘stop-and-think’

responses to events. Except that, unlike electric shocks, metacognitive feelings occur where there is an unexpected lack of fluency in mental processing—and so often enough, there is potential for learning. But in isolation, a metacognitive feeling does not tell you anything much about what could be learnt.

I mentioned earlier (in Section 7.2) Koriat’s proposal that ‘metacognitive feelings ... allow a transition from the implicit-automatic mode to the explicit-controlled mode of operation’ (Koriat 2000, 150). Although he wasn’t talking about development, I think we can productively misinterpret what he says as about development. Metacognitive feelings have a dual role. By triggering ‘stop-and-think’ responses to events which interfere with automatic processing, they may create opportunities for learning. But because they are intentional isolators, they also serve to keep the later-developing, less automatic processes separate from the more automatic, early-developing processes. So they both ‘allow a transition’ of one kind and prevent a transition of another kind.

The challenge we face, then, is to explain how rediscovery might occur. Coming to know simple facts about particular physical objects may begin with object indexes and the metacognitive feelings these give rise to, but it does not end there. Coming to know simple facts about physical objects may involve interacting with others around you who already have knowledge. This is one reason for investigating infants’ abilities to track others’ minds and actions, and to act together with others.

Interlude on innateness

9 Innateness

If you are going to discuss whether something is innate, you had better fix on a particular notion of innateness. There are plenty to choose from, although no one notion seems to be entirely satisfactory (as Mameli and Bateson 2011 argue). I will simply stipulate that for a cognitive ability to be innate is for its developmental emergence not to be a direct consequence of data-driven learning. In short, innate = not learned.

This way of characterizing innateness fits perfectly with the most compelling arguments for innateness, which take the form of poverty of stimulus arguments. Such arguments aim to show that some aspect of humans’ syntactic abilities, say, are not entirely a consequence of data-driven learning (see Pullum and Scholz 2002, for details). As we will see in a moment, there are some compelling poverty of stimulus arguments. So any acceptable characterization of innateness must respect the fact that poverty of stimulus arguments can establish innateness. And equating innate with not learned, as I propose, is the simplest way to achieve this.

Samuels objects to equating innate with not learned (see Samuels 2004, 139). He notes that it is unilluminating, on the grounds that the notion of learning is difficult to characterize for much the same reasons that the notion of innateness is. Samuels aims for a deeper understanding of innateness. I think Samuels is right that equating innate with not learned is barely informative, but I see this as a virtue rather than a deficit. If the last couple of millennia or so are any guide, philosophical methods do not enable us to give substantial characterizations of non-logical phenomena. With this in mind, modesty in philosophical characterization can be a virtue—we aim merely for theoretical coherence, recognizing that discovering substantial truths usually requires doing more than just thinking about them. This is why I characterize innateness simply as a matter of not being a direct consequence of data-driven learning.

9.1 Syntax

The domain in which innateness has been most extensively and carefully researched is probably syntax. What is syntax? Contrast these two sequences (to adapt a famous example from Chomsky):

  1. The turnip of shapely knowing isn’t yet buttressed by death.
  2. *The buttressed turnip shapely knowing yet isn’t of by death.

Whereas the second sequence of words is not a sentence, the first is widely recognized as a sentence—although not one you are likely to make much sense of. Now take a second contrast:

  1. Ayesha ate Ben.
  2. Ben ate Ayesha.

Again, the two sentences involve the same words, but they have quite different implications. The difference is syntax.

Although capable of tracking syntactic differences and exploiting them in communicating with words, humans are often unaware of syntax. Consider a simple phrase, ‘the red ball’. In principle, this phrase could have two different syntactic structures, as illustrated in Figure 9.1. The difference in structure may not appear very significant at first. But consider what would happen if someone said, ‘This red ball is broken, please bring me another one!’ Can you comply by bringing a blue ball, or does the ball have to be red? The answer turns out to depend on the structure of the phrase ‘this red ball’. If you think only a red ball will do, you are treating the phrase as having nested structure (see Figure 9.1). This is because the ‘one’ in ‘please bring me another one’ can only refer to something that a constituent of an earlier part of the sentence already introduced, and if ‘this red ball’ has a flat structure, ‘red ball’ is

Figure 9.1 Two possible structures for ‘the red ball’

Source: Lidz, Waxman and Freedman (2003).

not a constituent. In advance of having this explained, many people (me included) would have no idea which structure, flat or nested, they assign to phrases like ‘the red ball’.

Abilities to track syntactic properties and exploit them in communicating with words cannot be based merely on past exposure to particular sentences. After all, you can often distinguish entirely novel sentences from non-sentences. To describe the ability, it seems that we will need principles of syntax, much as we needed principles to characterize abilities to segment objects, represent them as persisting and track their causal interactions (see Chapter 2). We can think of the syntactic principles as entailing, for any arbitrary sequence of words, whether or not it is a sentence.

9.2 A poverty of stimulus argument

How can we discover whether syntactic abilities are innate? Lidz, Waxman and Freedman (2003) set out to answer this question using phrases like ‘the red ball’. As we saw in Section 9.1, such phrases can in principle be assigned either of two syntactic structures, one flat and the other nested. (These were illustrated in Figure 9.1.) Lidz, Waxman and Freedman (2003) aimed to show that 18-month-old infants already interpret this sort of phrase as having a nested structure, and that they could not have learnt to do this just on the basis of their experiences of language. This would imply that at least some aspects of humans’ syntactic abilities are a consequence of things that are innate.

The first task is to determine how 18-month-old infants interpret phrases like ‘the red ball’. Suppose someone says, ‘Look, a red ball! Do you see another one?’ What would you be looking for? If you were to look for any ball, whether red or not, this would indicate that you had interpreted the phrase as having a flat structure. But if you are like me, you will be looking for another red ball. This indicates that you interpret the phrase ‘the red ball’ as having a nested structure. We can therefore find out whether some people interpret the phrase as having a nested or a flat structure by measuring whether they would look for a red ball or just any ball. This is the perfect approach for 18-month-old infants, since it does not require them to make any verbal responses. Accordingly, Lidz, Waxman and Freedman (2003) played infants a recording of sentences like ‘Look, a red ball! Do you see another one?’ While they heard this recording, they could see a red ball (or other object corresponding to the sentence they heard). After the recording, the infants were shown two things simultaneously, such as a red ball and a blue ball. It turned out that infants looked significantly longer at the red ball than the blue one.

You might object that this finding is not revealing. After all, infants’ tendency to look longer at the red ball than at the blue ball might be due to the fact that they have just seen a red ball. To rule this out, Lidz, Waxman and Freedman had a control condition. This was exactly like the main condition except that the sentence infants heard was like ‘Look, a red ball! What do you see now?’. That is, instead of asking ‘Do you see another one?’, they asked ‘What do you see now?’. In this control condition, infants showed the opposite pattern in their looking times: they looked significantly longer at the blue ball or other novel object (see Figure 9.2). So infants’ tendency to look longer at the red ball really does indicate that they interpret this phrase as having a nested structure.

Figure 9.2 In response to ‘Look, a red ball! What do you see now?’ (‘Control’), infants look longer at the blue (‘novel’) ball. But in response to ‘Look, a red ball! Do you see another one?’ (‘Anaphoric’), infants look longer at the red (‘familiar’) ball.

Source: Lidz and Waxman (2004), Figure 1.

The fact that 18-month-olds tend to interpret phrases like ‘the red ball’ as having a nested structure is not by itself evidence of innateness. We also need evidence that infants could not have learnt the nested structure on the basis of experiences of language. Lidz, Waxman and Freedman (2003) aimed to provide this by analysing a large collection of around 45,000 utterances directed to children and infants. In all of these utterances, they found only two which could have indicated nested structure to infants. To put this into perspective, there were four ungrammatical uses of ‘one’. This survey suggests that 18-month-olds’ experiences of language do not provide them with a basis for learning that phrases like ‘the red ball’ have a nested structure.

Putting Lidz, Waxman and Freedman’s (2003) two findings together gives us a poverty of stimulus argument. The two findings are that 18-month-olds can interpret phrases like ‘the red ball’ as having a nested structure, although their experiences of language do not provide them with a basis for learning this. The poverty of stimulus argument goes like this:

  1. 18-month-olds can interpret phrases like ‘the red ball’ as having a nested structure.
  2. To acquire this ability by data-driven learning would require experiences of utterances in which such phrases had to be understood as having a nested structure.
  3. But extremely few such utterances are directed to 18-month-olds.
  4. So 18-month-olds do not acquire the ability to interpret nested structure by data-driven learning.
  5. But all acquisition is either data-driven learning or innately primed.
  6. So 18-month-olds’ acquisition of the ability to interpret nested structure is innately primed.501

Note that this argument establishes that something is innate but does not tell us what is innate. The conclusion is not that the ability to interpret phrases like ‘the red ball’ as having a

nested structure is innate. It is that acquiring this ability depends on something that is innate. Establishing what that innate thing is would require further work.

Some linguists assume that some specifically syntactic abilities must be innate. For instance, Chomsky (1965, 25) describes the linguists’ task as that of characterizing the ‘innate linguistic theory that provides the basis for language learning’. Poverty of stimulus arguments provide no justification for this assumption. Further evidence would be needed to support a conclusion about what in particular is innate.

How strong is Lidz, Waxman and Freedman’s (2003) argument for innateness? The first premise is that 18-month-olds can interpret phrases like ‘the red ball’ as having a nested structure. To date, this depends on evidence from a single lab, so should be regarded with caution. The third premise is that few utterances in which phrases like ‘the red ball’ have to be interpreted as having a nested structure are directed to infants. Confidence in this premise should be high, given that it is based on a large number of recorded utterances. However, there is a possible line of objection linked to premise 2, which is about the need for such utterances. Lidz, Waxman and Freedman (2003) considered utterances directed to infants where the nested interpretation of phrases like ‘the red ball’ is required to understand the sentence. But it is possible in principle that infants may be exposed to some combination of linguistic and non-linguistic evidence that enables them to arrive at the nested structure interpretation. Attempts to develop a concrete objection along these lines (see Akhtar et al. 2004) are not convincing (Lidz and Waxman 2004, 161–2). While new discoveries about evidence or learning mechanisms are always possible, as things stand, the balance of evidence appears to favour the view that some syntactic abilities depend on something innate.

The discovery that something innate underpins some of infants’ syntactic abilities is a major breakthrough. It shows that there is no general reason to hold that development in other domains could not depend on things that are innate. And reflection on the case of syntax supports the view that early-developing abilities and states, perhaps including core knowledge, play an important role in development. This is a major challenge to theories about the developmental emergence of knowledge which exclusively invoke social interaction.

9.3 The poverty of poverty of stimulus arguments

Innateness is an exciting, attention-grabbing topic. Why risk destroying interest in it by focussing so narrowly on the case of ‘the red ball’? Consider an example of how poverty of stimulus arguments have been wielded in philosophy: ‘There would seem not to be enough ambient information available to account for the functional architecture that minds are found to have’ (Fodor 1983, 35). It is hard to detect an argument here. At least, if there is an argument, an equally compelling argument can be obtained by deleting the word ‘not’ from this sentence.

Things are not very different in linguistics. In a recent defence of poverty of stimulus arguments for a conclusion about an innate basis for syntactic abilities, Berwick et al. (2011) cite no evidence at all concerning the experiences available in development. They also cite no evidence at all concerning the development of syntactic abilities. When making this sort of observation, I am usually told that the evidence is of a general nature and too familiar to

need explicit mention. But this is not quite right. Instead Berwick et al.’s (2011) argument has a familiar form, which I propose to label the poverty of theory argument:

  1. Current theories about how certain syntactic abilities are acquired are inadequate.
  2. So the acquisition of such abilities depends on innate representations of syntactic structure.

It is important to distinguish poverty of theory arguments from poverty of stimulus arguments. Poverty of theory arguments could be used to establish that almost everything is innate. After all, fully adequate developmental theories are scarce. The problem, of course, is that poverty of theory arguments assume that the inadequacy of theories is not due to the inadequacy of theorists. By contrast, poverty of stimulus arguments do not require this assumption. Poverty of stimulus arguments therefore provide potentially more convincing grounds for accepting conclusions about innateness.

You might think that poverty of stimulus arguments have been around for a long time and are widely used to establish innateness, especially concerning syntax. But it turns out that this is a myth. In a thorough review, Pullum and Scholz (2002) observe that ‘the APS [poverty of stimulus argument] still awaits even a single good supporting example’ (2002, 47). Shortly after they wrote this, Lidz, Waxman and Freedman (2003) published their example involving ‘the red ball’. And that is the best, most careful attempt to provide a poverty of stimulus argument for human syntactic abilities to date.

9.4 Is core knowledge innate?

Core knowledge is often characterized by a list of properties including innateness (for example, Carey and Spelke 1996, 520). Why accept that all, or even any, core knowledge is innate?

Poverty of stimulus arguments reveal that some abilities are innate in non-humans (for example, Chiandetti and Vallortigara 2011), and perhaps that something underpinning the acquisition of some syntactic abilities is innate (as we have just seen). So there is no general reason to oppose claims about innateness. On the other hand, in no domain other than syntax has a poverty of stimulus argument been provided for humans. (And even in the case of syntax, the critical evidence is not overwhelming.)

Why are even some of the most careful researchers willing to assert that core knowledge is innate? Sometimes the arguments appear to be variants on ‘poverty of theory’ arguments. For example: ‘It is far from clear how children could learn anything about the entities in a domain ... if they could not single out those entities in their surroundings’ (Spelke 1994, 439). But there are also more interesting lines of argument:

If early knowledge encompasses environmental constraints that are not obvious in the child’s perceptual and motor experience while failing to encompass more obvious constraints, then this knowledge is not likely to have been shaped by the child’s perceptual and motor experience.

(Spelke 1994, 438)

We might call this an irrelevance of the stimulus argument. If the stimuli (or data) available to infants are unconnected to an ability they are acquiring, we might suppose that the ability is not entirely a consequence of data-driven learning. While an interesting idea, this idea seems to have limited force. One is that far more information is processed perceptually and motorically than is experienced; and experience is not surely necessary for learning to take place. The other is a weakness shared with informal poverty of stimulus arguments: in both cases, the informal arguments rest on unargued claims about the availability of information and what learning processes might extract from it.

In the absence of evidence for the innateness of core knowledge, and given how little is currently known about developing minds, it seems to me that we should be agnostic. We do not currently know whether any core knowledge is innate.502

9.5 Syntax and rediscovery

This is really a chapter about innateness; but discussing innateness requires reference to syntax, which provides a perfect opportunity to highlight the idea that the developmental emergence of knowledge is a process of rediscovery.

To make the connection, we should first return to the three-fold distinction between formally, descriptively, and explanatorily adequate (see Table 3.2). Consider formal adequacy. Imagine someone who is omniscient except concerning which sequences of words are sentences, and who has unlimited cognitive resources. Suppose this person took certain principles of syntax to be true concerning a particular language. Could she now say, for any sequence of words, whether it is a sentence of that language? If she could, these principles are formally adequate for the language. Although constructing formally adequate principles of syntax for a language turns out to be a difficult problem (for example, Chomsky 1957, 13–25), it seems clear that the problem can be solved (see Jackendoff 2003a, Chapter 3, for an introduction to one approach). It is also plausible that we could identify principles of syntax which are descriptively adequate for a particular group of speakers. These principles would enable us to predict which sequences of words this group of speakers will identify as sentences. The controversial issue is explanatory adequacy. Just here we run up against a counterpart of the Linking Problem (from Section 4.4). What is the relation between principles of syntax that are descriptively accurate for a particular language user and the mechanisms which underpin her syntactic abilities?

Inspired by the Simple View (from Section 3.1), we might start with the idea that the language user knows principles of syntax. What links the principles to an individual thinker is the thinker's knowledge of the principles. This enables her to apply the principles in just the way a scientist would: she can make inferences from them to determine whether a given sequence of words is a sentence or not. The problem is that adults are typically ignorant of syntax. Worse, they are susceptible to false beliefs which do not directly impair their performance. So the counterpart of the Simple View for syntax is uncontroversially false.503

Since we do not know what states might link descriptively adequate principles of syntax to an individual thinker, we might as well call the things that provide this link, whatever they turn out to be, her 'core knowledge' of syntax.504

How does knowledge of the syntactic aspects of utterances emerge in development? It is usually not until relatively late in life, and often as a consequence of formal education, that humans ever come to know facts about syntax, if they do at all. As Fodor notes, in coming to know these facts, we are rediscovering something that is in some sense already implicit in core knowledge of syntax: 'when you learn about English syntax (e.g., in a linguistics course), what you are learning is something that, in some sense, you already knew' (Fodor 1983, 134, footnote 23).

Further, the process of coming to know facts about syntax is clearly not a matter of transforming the contents of core knowledge states (if indeed they have contents) into the contents of knowledge states. That is, the Assumption of Representational Connections would not be correct in the case of syntax. Core knowledge does not later become knowledge proper. Instead, gaining knowledge of syntax is a paradigm case of development as rediscovery.

What does rediscovery mean? The things to be known are in some sense already implicit in your core knowledge, or in early-developing capacities in you, such as the capacities which enable you to learn about the nested structure of phrases like 'the red ball'. Yet when you come to acquire the knowledge, you do so in roughly the way a linguist does. The basis for learning is not the early-developing syntactic abilities but the things they enable you to do: to produce utterances which in fact have certain syntactic features, and to discriminate among utterances according to what are in fact their syntactic features. Reflection on the behaviours and patterns of discrimination on your part, together with experience of these and social interaction with others typically involving instruction, eventually enables you to acquire knowledge about syntax (see Figure 9.3).

As in other cases, so for syntax: core knowledge or other early-developing abilities may play a key role in the emergence of knowledge not by providing building blocks for the contents of knowledge states but by enabling you to do, feel and experience things, and reflection on these, together with social interaction, plays a role in facilitating the emergence in the development of knowledge.

Figure 9.3 The developmental emergence of knowledge of syntax involves rediscovery

Figure 9.3

9.6 Conclusion

To be innate is to emerge in development otherwise than as a direct consequence of data-driven learning. The only known way to establish that something is innate is by way of a poverty of stimulus argument. There has been striking success in identifying evidence for innateness in a single case, 'the red ball' (see Section 9.2), and in comparative research (for example, Chiandetti and Vallortigara 2011). These successes allow us to conclude that early-developing abilities can depend on things which are innate.

But contrary to any impression that developmental scientists are deeply engaged in battles between rationalism and empiricism, progress in providing genuine poverty of stimulus arguments concerning humans has been limited to one case. Given the dramatic progress that has been made in understanding other aspects of development, this indicates that debates about innateness might not be very important after all.505 Certainly we should be agnostic on whether core knowledge in domains other than syntax is innate; and even in the case of syntax, we should recognize that very little is known about what is innate.

Notes

Part II

Minds and actions

10 Action

There are two fundamental ways of specifying an action. Suppose you reach out and grasp something nearby, a mug, say. Your action can be specified in terms of the joint displacements and bodily configurations involved. For example, your elbow joint becomes straighter as you reach out, your fingers pre-shape the mug or its handle and your wrist rotates. Alternatively, your action can be specified in terms of its goal, namely, to reach out and grasp the mug. Importantly, actions involving quite different patterns of joint displacements and bodily configurations can all be actions directed to this goal. Change even a tiny detail about the mug you were grasping by, for example, by rotating the handle away from you, and you will systematically alter the joint displacements and bodily configurations involved without changing the goal.

Knowledge of goals is essential for understanding others' thoughts and actions. I seize little Isabel by the wrists and swing her around, thereby making her laugh and breaking a vase. Looking on, you might wonder what the goal of my action was. Did I act in order to break the vase or just to make Isabel laugh? Or was my action perhaps directed to some other goal, one not realized because my action failed? Knowing facts about the goals of others' current and imminent actions is every bit as essential for social animals like humans as knowing facts about the movements and interactions of merely physical objects. But how do humans first come to have such knowledge?

10.1 Tracking vs knowing

Start by fixing terminology. Among all of the actual and possible outcomes of a purposive action, some are outcomes to which the action is directed. These outcomes are the goals of that action. Note that goals are not intentions, nor mental states of any kind. Take someone who tells you that the goal of her actions is the release of a political prisoner. She is not

talking about her intentions: she's talking about a (currently nonactual, but possible) state of the world. This is potentially confusing because others sometimes use the term 'goal' to refer to mental states in virtue of which actions are directed to outcomes. But I will always use 'goal' to refer outcomes to which actions are directed.

To know simple facts about the goals of an action is to know something about to which outcomes that action is directed. By contrast, to track the goals of an action is merely for there to be a process in you whose unfolding non-accidentally depends in some way on which outcomes are the goals of that action. It is possible, in theory at least, to track the goals of actions without knowing or representing any facts about goals at all.

To illustrate, consider an analogy. An ant that has exploited a food source will lay a pheromone trail on its way back to the colony. Other ants may then detect this pheromone trail, follow it and exploit the food source for themselves (Sumpter and Beekman 2003). These ants are tracking the locations of food in their environment. But their doing so does not involve knowing about or representing locations, nor does it require a model of space. Instead what the ants identify are the pheromone trails. So identifying pheromone trails enables ants to track the locations of food sources without knowledge of locations. As this illustrates, tracking something does not always involve knowing about it, nor even representing it.601

In this chapter we will first consider research showing that infants, from 3 months of age or earlier, can track the goals of actions. Whether this tracking involves knowledge of goals is a further question we will come to later.

10.2 Three-month-olds track the goals of actions

Gergely et al. (1995) asked what happens when 12-month-olds see an action. Do they represent the movements involved in the action only, or are they also sensitive to the goals of the action? To answer this question, Gergely et al. (1995) compared two changes: in one, there is a change in movement trajectory but no change in goal; in the other there is a change in goal but no change in movement trajectory. If infants track trajectories only, they should be more interested in the former change. But what Gergely et al. actually found was that their 12-month-olds were more interested when the goal changed even though this involved no change in movement trajectory.

How did their experiment work? Gergely et al. habituated all the infants to the animation represented in Figure 10.1. There are two balls in this animation, a small one and a large one. The animation starts with cues suggesting affinity between the two balls and indicating that the small ball's movements are self-propelled. The small ball then moves over the barrier and stops by the large ball. Once infants were habituated to this animation, they were shown one of the two new animations represented in Figure 10.2. In the 'New Action' condition, infants saw the small ball move directly to the larger ball, whereas in the 'Old Action' condition the small ball follows the same trajectory it moved along when there was a barrier between it and the large ball. Thinking purely in terms of objects' trajectories, the animation shown in the New Action condition is most different from the animation to which infants were habituated. So if the infants track trajectories only, they should show greater dishabituation in the New Action condition than in the Old

Figure 10.1 The small ball moves over the barrier and stops by the larger ball. This is an illustration of part of a movie used in an experiment on infants' abilities to track the goals of actions.

Source: Gergely et al. (1995), Figure 1(b).

(a) New Action

(b) Old Action

Figure 10.2 Following habituation, infants were shown one of the two movies represented above. In (a), the small ball moves directly to the large ball; in (b) the small ball takes the same trajectory taken when there was a barrier between it and the larger ball.

Source: Gergely et al. (1995), Figure 3.

Action condition. By contrast, the hypothesis that the infants can track goals to which actions are directed is consistent with the opposite pattern of dishabituation. To see why, imagine you had seen the initial animations and concluded that the small ball's actions are directed to the goal of reaching the large ball. Then you are in the Old Action condition and see the small ball following the old trajectory. Seeing this should suggest that you have missed something—if the goal of the small ball's actions is merely to reach the large ball, it is unclear why it is taking such an indirect route to the large ball. And Gergely et al. did indeed find that the Old Action condition produced greater dishabituation in their infants. They concluded that 'by the end of the first year infants are … capable of … interpreting the goal-directed behavior of rational agents' (1995, 184).

You might object that it is bizarre to use balls in a study about actions. After all, it seems unlikely that many 12-month-olds really think that mere balls perform goal-directed actions. But it turns out that human adults are surprisingly willing to detect goals (and even motives and needs) in patterns of movement involving geometric shapes (Heider and Simmel 1944). It is a reasonable guess that infants, like adults, apply whatever capacities they have for tracking goals to mere balls (although there may be grounds to doubt this guess, as we will see in Section 11.4). And of course there is an excellent reason for Gergely et al. (1995) to use balls rather than more realistic depictions of agents. The simplicity of their animations enabled them to create carefully matched control conditions, and so to exclude with more confidence the possibility that their results were an artefact of some irrelevant feature of the animations. But is such simplicity necessary or are infants in the first year of life also sensitive to the goals of actions performed by ordinary humans?

Evidence that they are comes from an elegant and much-replicated study by Woodward (1998, study 3). In Woodward's paradigm, infants are presented with a scenario involving two objects, a ball and a teddy, say. In the habituation phase, a hand enters the scene and grasps one of the two objects, as depicted in Figure 10.3. In the main phase, infants are then

Figure 10.3 Schematic representation of habituation and test events from an experiment by Woodward (1998) on 6-month-olds' abilities to track the goals of actions

Redrawn from Woodward, Sommerville and Guajardo (2001), Figure 1.

Figure 10.3

shown one of two new events. As in Gergely et al.'s (1995) experiment, the idea is to pit sameness of movement trajectory against sameness of goal. In both new events, the locations of the ball and teddy are reversed. In one event, the hand follows the same trajectory as before and so grasps a new object. In the other event, the hand grasps the same object as before, which requires following a new trajectory. If infants ignore goals and track trajectories only, then they should find the event with the new trajectory more interesting. But if infants track the goals of actions, then they may well be more interested in changes in the goals of actions than in changes in movement trajectories. In that case, they should find the event with the new goal more interesting and so dishabituate more strongly to it. And this is just what Woodward observed in both 9- and 6-month-olds. She concluded that 'early in life, infants begin to set up a system of knowledge of human action that has features in common with more mature understandings, and that is distinct from their knowledge of inanimate object motion' (Woodward 1998, 31).

When can humans first track the goals of purposive actions? Using an ingenious manipulation that we shall discuss later ('sticky mittens'; see Section 10.7), Sommerville, Woodward and Needham (2005) created a variation on Woodward's paradigm to show that even 3-month-olds can form expectations based on the goal of an action (for another study with 3-month-olds, see Luo 2011; and Gredebäck and Melinder 2011 on 4-month-olds). This makes sense. Humans are social animals and tracking the goals of actions is necessary for almost any kind of social cognition. So we might expect that they can track the goals of actions as early as they can track the merely physical behaviours of objects.

10.3 Pure goal tracking

We have seen that infants, even at 3 months of age, can track not only movements but goals to which actions are directed. How do they do this? Tracking goals is an astonishingly complex problem unless you can communicate with others or already know something about their intentions. You start with information about joint displacements and bodily configurations, and, if things go well, you somehow end up with an outcome to which those all are directed. How is this possible?

In tracking the goals of some actions, it may be that adults are ascribing intentions and other mental states to the agent of those actions. Someone watching you feed the infant may be thinking that you are acting on an intention to feed it, for example. Are infants also ascribing intentions when tracking goals?

Many suggest that they are. For instance, Woodward (2009, 53) suggests that 'infants understand intentions as existing independently of particular concrete actions and as residing within the individual', and she remarks that this 'is essential to recovering intentions from observed actions' (see further, Premack 1990, 14; Woodward, Sommerville and Guajardo 2001, 168).

In considering these claims it is vital to separate two issues. One issue is whether infants in the first year of life ever track agents' intentions as well as the goals of their actions. The other issue is whether infants (and adults) can track goals independently of any knowledge of, or information about, mental states.

On the first issue, as far as I know, there are as yet no experiments which directly address this issue (which means not confounding tracking intentions with tracking goals).602 Research on infants' understanding of other mental states has uncovered plenty of evidence that infants in the first year of life can track at least some mental states (see Chapter 12). This should make us at least open to the as-yet-untested possibility that these infants can track intentions too. But this is not an issue that will concern us in this chapter, as our aim to gain a more fundamental understanding of infants' abilities to track the goals of actions.

The second issue is whether infants and adults can track goals without relying on any information about mental states. To see why this issue matters, note that even simple actions typically involve a hierarchy of goals. To illustrate, consider feeding an infant. You reach for, grasp and pick up the spoon, scoop up some goo and deftly pilot the spoon around the infant's flailing arms and towards her opening mouth, spontaneously opening your own mouth at the same time as if to show the infant what to do. Considered as a whole, the goal of your action might have been to feed the infant. But many components of your action also have goals in their own right. Reaching for the spoon, grasping it and picking it up are all goals relative to which you succeeded but could have failed. These smaller goals are related to the larger goal of feeding the infant as means to that end.

Now imagine you are an infant being spoon-fed your first solids by an adult. To track the goal of the adult's action when feeding you, you need to identify the smaller goals and not only the larger goal. After all, achieving any of these smaller goals, such as reaching for the spoon, can involve indefinitely many different patterns of joint displacements and bodily configurations. As the adult repeats the larger action, putting one spoon after another into the infant's mouth, the joint displacements and bodily configurations vary enormously while the goals are relatively invariant. It would be inefficient and probably infeasible to move directly from patterns of joint displacements and bodily configurations directly to the goal of feeding the infant. Instead, goal tracking starts with very small bits of purposive action like reaching, grasping and transporting.

Assume (for the sake of argument) that you, as the infant being fed solids for the first time, know what intentions are and can ascribe them. Still, on what basis can you determine the intentions behind the adult's actions? You can't communicate linguistically with them. In fact, it seems that the only access you have to another's intentions is via the actions they perform. Suppose that to track the goals of actions, you have to identify the intentions with which those actions are performed. This is a significant burden. What are the intentions of the adult who has just grasped the spoon? If spoon feeding is completely new to you, you are in no position to detect that she intends to feed you. You will therefore have to identify intentions with which smaller actions are performed, like the intention of grasping the spoon. But consider what this involves. To be tracking intentions and not merely goals, you need to be sensitive to the possibility that intentions and goals can come apart. (Otherwise there would be insufficient reason to suppose that you are tracking intentions.) You need to be sensitive to the possibility that the action she performs in grasping the spoon is directed to the goal of grasping the spoon although she had no intention to grasp the spoon at all—perhaps, for example, her intention was to grasp the fork, but she got distracted and acted counter to he own intentions. Detecting that intentions and goals have come apart requires relating

communicative error signals to particular actions ('oops'), or tracking relations between the goals of very small actions and the goals of larger sequences of actions. So if all goal tracking requires requiring information about mental states, it depends on a relatively rich appreciation of the structures of sequences of actions.

This is why the second issue matters: if all goal tracking requires information about mental states, the complexity of the structures infants are tracking means that it will be very difficult to understand how this is possible, even in principle. By contrast, suppose some goal tracking can be done independently of any information about mental states. Then not only is the task of constructing a theory of goal tracking significantly simpler. Also, it is theoretically coherent to conjecture that infants (and adults) would be able to first track the goals of very small actions and learn about the structure of larger sequences of actions, so as eventually to use information in ascribing intentions to agents. Indeed, in this case, goal tracking could be a foundation for the ascription of mental states and the identification of expressions of emotion, for social interactions and for communication.

So can goal tracking in infants (or adults) occur independently of tracking intentions and other mental states? Let us say that pure goal tracking is goal tracking which does not involve ascribing intentions or any other mental states. Woodward and others who assert that infants understand intentions may do so because they assume that ascribing intentions is an essential part of goal tracking, which would imply that pure goal tracking is impossible, even in theory. But is pure goal tracking theoretically impossible? And if not, does it actually occur in infants or in adults?

10.4 The Teleological Stance

The fact that infants can track the goals of actions from 3 months of age or earlier (see Section 10.2) motivates us to ask a question. How do they do this? Section 10.3 argued for the value of considering the possibility of goal tracking that is independent of any information about mental states, that is, of pure goal tracking. Now we would like to know how in principle an infant (or anyone) might track the goal of an action.

Csibra and Gergely (1998) have argued that pure goal tracking is possible and can be achieved by assuming what they call the 'principle of rational action':

For an outcome to be among the goals of an action is for the joint displacements and bodily configurations which realise that action to be the best way to achieve that outcome that is available to the agent given current constraints on her possibilities for action.603

How does this work? To illustrate, suppose you are observing someone who is moving in the way represented in Figure 10.4. You might wonder whether the goal of her action is to reach the blue house (House A) or whether it is to reach the red house (House B). Since her movements are not the best available way to reach the blue house (House A). we can use the above principle to exclude the possibility that her goal is to reach the blue house (House A). In general, the strategy is this. Start with a set of outcomes that are candidate goals and, for

Figure 10.4 Is the goal of these movements to reach House A or to reach House B?

Figure 10.4

each outcome, ask yourself whether the observed joint displacements and bodily configurations are the best way to achieve that outcome available to the agent given the current constraints. Whenever the answer is no, exclude the outcome. Any remaining outcomes can be regarded, at least provisionally, as among the goals of the agent's actions.

In applying this principle, we are adopting what Csibra and Gergely call the Teleological Stance. They do not claim that the principle is true (which is good, as the principle is clearly false). Instead their claim is that assuming the principle would enable us, in a limited but useful range of situations, to accurately track the goals of actions. If this is correct, they have demonstrated that pure goal ascription is possible.

One limit on the Teleological Stance as so far characterized is that it could not underpin abilities to track the goals of actions that fail. To illustrate, consider someone who, intending to chop zucchini, accidentally slices a sliver of skin off the end of her own finger. Of course the goal of her action was to chop the zucchini rather than her finger. But as so far characterized, the Teleological Stance implies that the opposite was true. For since a better way to chop the zucchini was available to her given the current constraints, the above principle implies that chopping the zucchini could not have been a goal of her actions (which, by stipulation, it was). Further, since no better way of taking off the end of her finger was available to her, the above principle will never enable us to exclude the hypothesis that doing this was among the goals of her action (which, by stipulation, it was not). As this illustrates, the Teleological Stance seems to be limited to actions that succeed.604

The existence of this limit is not in itself an objection to any claim we have yet considered. But abilities to track the goals of actions which fail are important, and clearly present in infancy at least by 9 months of age (Behne et al. 2005). Can we refine the Teleological Stance to

overcome this limit? As a first step, we might add a second principle to Csibra and Gergely's 'principle of rationality', one which must also be satisfied: 'Outcomes which are among the goals of an action are outcomes bringing about which would typically be desirable to animals like the agent of the action.' This excludes taking off the end of her finger as a goal of the failed zucchini-slicer's actions, as bringing about this outcome is not typically desirable. But can we invoke desirability in elaborating a theory of what is supposed be pure goal tracking? Wouldn't doing so violate the requirement that pure goal tracking cannot involve ascribing mental states? In fact, it would not. To see why, note that there are objective notions of desirability. Ayesha believes herself stuck in London for the night and has no inkling that she could still catch the last train because its departure has been greatly delayed. Although she therefore does not want or intend to go to the station, we as outside observers might comment that it would be desirable for Ayesha to go to the station. In characterizing pure goal tracking, we can safely appeal to a non-mentalistic notion of desirability such as this in adding the above principle.

Appealing to desirability is not quite sufficient to overcome the limit on tracking the goals of actions which fail. Considering desirability will often be sufficient to exclude the actual outcome of a failed action as a goal of that action. But actions which fail are not usually the best available ways to achieve the outcomes which are actually their goals. In the zucchini case, for example, moving the finger further from the blade's trajectory would be a better way to chop the zucchini. For this reason, it appears that the Teleological Stance as so far characterized could not underpin successful pure goal tracking when actions fail.

To overcome this limit, we do not need to add or refine the two principles already suggested. Instead we need to consider the relation between tracker and trackee, that is, between the individual who is tracking goals and the agent of the actions being tracked. The trackee's task is to select the best available way to achieve the outcomes which are in fact the goals of her action. In effect, she is computing joint displacements and bodily configurations given goals and constraints on her possibilities for action. The tracker's task is to discern which outcomes the observed joint displacements and bodily configurations are the best way of achieving given the current constraints on the agent's actions. In effect, she is computing goals, given joint displacements and bodily configurations and constraints on action. For the purposes of goal tracking, it is not important that the tracker be especially good at identifying the best available ways to bring outcomes about. To maximize the chances of successful goal tracking, what matters is that the tracker and trackee are as similar as possible. When the trackee misses the zucchini, the tracker can still successfully track the goals of her actions providing she relies on similarly flawed processes to compute the best available ways of achieving outcomes. As we will see (in Section 11.2), there is a way of implementing the Teleological Stance which is well suited to exploiting similarities between the trackee and the tracker.

It may be possible to further refine the Teleological Stance. However that is done, some limits on the range of situations in which it enables successful goal tracking will surely remain. Whatever its limits eventually turn out to be, Csibra and Gergely's discovery of the Teleological Stance is exciting because it demonstrates the theoretical possibility of pure goal tracking.

But what might the Teleological Stance tell us about infants' goal tracking? The Teleological Stance can be thought of as a principle characterizing action, much as the Principles of Object

Table 10.1 Three questions about the Teleological Stance

| Formal Adequacy | If someone took the Teleological Stance to be true, was omniscient about joint displacements and bodily configurations, and had unlimited cognitive resources, to what extent would she be able to track the goals of actions? | | Descriptive Adequacy | Does the Teleological Stance enable us to generate correct predictions about infants' and others' abilities to track the goals of actions? | | Explanatory Adequacy | Is there a link between the Teleological Stance and infants' or others' minds, and does this link partly explain how it is they are able to track the goals of actions? |

Perception characterize objects and their interactions (see Section 2.3). In both cases, we can distinguish three questions (see Table 10.1). First, this section has been argued that the Teleological Stance is formally adequate: it explains the possibility in principle of pure goal tracking, or could be made to do so with further refinements. The next questions we should consider are whether the Teleological Stance is descriptively adequate and, if it is, whether it is also explanatorily adequate. Given how things went in Part I, you might be anticipating that our focus will be on the question of explanatory adequacy. This is correct, although in the next section we will see that the question of descriptive adequacy is not entirely straightforward.

10.5 Statistical regularities

Is the Teleological Stance descriptively adequate? Gergely et al. (1995) show that there are cases in which infants' goal tracking does indeed conform to the Teleological Stance (see Section 10.2), and there are many further studies extending this one (see Csibra 2003, for a review). But, confusingly, there are also cases which indicate that the Teleological Stance is not descriptively adequate.

To see why, step back and think about a burglar watching your house for an opportunity to break in and steal your things. In predicting your actions, she is likely to be interested not only in the goals of your actions but also in your routines. She may know that you habitually leave home at 11:30 without knowing what your further goals in doing so are. As this illustrates, statistical regularities can be useful in predicting actions even without any deep insight into the goals of those actions.

We might expect that infants, like burglars, will sometimes make use of statistical regularities in their attempts to understand and predict actions. Indeed, there is evidence that they do. Paulus et al. (2011) showed 9-month-olds and adults a sequence involving a cow who comes to a fork in the road. The cow takes the upper road at the fork. This is the longer route to its destination but also apparently the best choice, as the lower road is blocked (see Figure 10.5). Infants watched this sequence repeatedly until habituated to it, and adults saw it eight times. They were then shown a new scene in which the lower, shorter road is no longer blocked (see Figure 10.6). Paulus et al. (2011) wanted to know which road the infants and

Figure 10.5 The cow crosses the scene using the top path and the bottom path is broken

Source: Paulus et al. (2011), Figure 1B (part).

Figure 10.6 The cow has entered the scene from the left and is now behind the oval occluder. Both the longer and the shorter path are available. Where do subjects anticipate the cow will emerge? Black rectangles show regions of interest for anticipatory looking.

Source: Paulus et al. (2011), Figure 1C (part).

Figure 10.5

Figure 10.6

adults thought the cow would take. If their predictions were based only on information about the goal of the cow's actions, they might well anticipate that the cow would take the newly open shorter road. But if their predictions were based on statistical regularities, then they should expect the cow to take the longer road since it had repeatedly done so in the past. And this is what both infants and adults in fact expected on first seeing the new sequence, indicating that their predictions were indeed based on statistical regularities (for further evidence, see Gredebäck and Melinder 2010; Green et al. 2016).

But how did Paulus et al. (2011) measure expectations? They used anticipatory looking. The fork in the cow's road was covered by an oval occluder (as shown in Figure 10.5). We know that when a moving object disappears behind an occluder, infants and adults alike will often proactively gaze to the place where they expect it to reappear. Accordingly we can detect which road an infant or adult watching the sequence expects the cow to take by measuring where they are looking in anticipation of the cow's emergence from the oval occluder. Paulus et al. observed that nearly all infants and adults who looked in anticipation of the cow's emergence from the oval occluder looked towards the upper road that it had always previously travelled along. This suggests that 9-month-olds sometimes use statistical regularities in anticipating actions.

Are these findings evidence that the Teleological Stance is not descriptively adequate? Paulus et al. appear to draw this conclusion. They note that the 9-month-olds, but not the adults, continue to anticipate that the cow will take the longer path even on seeing the new sequence for the fourth time. This, they suggest, contradicts the view that 9-month-olds track the goals of purposive actions in accordance with the Teleological Stance.605

But drawing this conclusion is not obviously quite correct. To see why, recall the burglar watching your house for an opportunity to break in. She has partial information about your goals and partial information about regularities in your behaviour. When the two kinds of information point in different directions, there isn't obviously anything wrong in prioritizing information about regularities over information about goals. Nor would her doing so reveal that she cannot track goals using the Teleological Stance. Similarly, findings that infants, like burglars, will sometimes make use of statistical regularities in their attempts to understand and predict actions are not evidence that they lack a goal-tracking ability for which the Teleological Stance is descriptively adequate.

In fact, using statistical regularities, as the burglar does, depends on goal tracking. To see this, note that the regularities in question do not concern patterns of joint displacements and bodily configurations but rather goal-directed actions. In leaving the house at 11:30, one day you wearing heels and the next day you are limping along in flats (following an unfortunate accident). The joint displacements and bodily configurations are markedly different: but what the burglar cares about is a regularity with respect to a goal, namely, the goal of leaving the house. Similarly, Woodward's experiment with the grasping hand (Woodward (1998) from Section 10.2) tacitly relies on the assumption that infants use statistical regularities. To generate the prediction that infants will find an action with a new goal more novel than an action with the same goal, it is not enough to suppose that infants can track the goals of actions: you must also suppose, further, that they can detect whether people do in future what they did in the past. But the regularity in question is specified in terms of a goal:

grasping (or touching, or reaching). So here, again, exploiting statistical regularities is not an alternative to goal tracking but something that builds on it.

So Paulus et al.'s (2011) findings should not convince us that the Teleological Stance is not descriptively adequate. But they do create an interesting problem. We cannot conclude that the Teleological Stance is descriptively adequate on the basis of those experiments in which infants' responses appear to fit this conclusion and also conclude that infants can make use of statistical regularities in goal tracking on the basis of other experiments. After all, this way of sorting the experiments all but guarantees that, whatever infants are actually doing, we will reach these conclusions. To avoid these conclusions being practically irrefutable, we need a richer theoretical framework, one that will enable us to understand why infants sometimes respond on the basis of statistical regularities (as in Paulus et al. 2011) and why they sometimes respond in line with the Teleological Stance (as in Gergely et al. 1995).

10.6 A methodological explanation?

We have just encountered a first puzzle about goal tracking in development (in Section 10.5). Why do infants sometimes respond on the basis of statistical regularities and sometimes respond in line with the Teleological Stance?

Daum et al. (2012) took a step towards addressing this question by adapting Woodward's paradigm with the hand reaching for and grasping a teddy or a ball (see Section 10.2 and Figure 10.3). The changes meant that Daum et al. (2012) could measure both anticipatory looking and dishabituation in a single individual observing a single scenario. Replicating Woodward's findings (see Section 10.2), patterns of dishabituation in 9-month-olds indicated that they use information about goals in forming expectations about actions. However the 9-month-olds' anticipatory looking indicated that they were relying more on statistical regularities in forming expectations about actions. Relatedly, Gredebäck and Melinder (2010) report a dissociation between failure to provide evidence of goal tracking in anticipatory looking but success in doing so when the measure was pupil dilation. (Pupil dilation indicates arousal and thus hints that an infant finds something unusual.) Why did patterns of dishabituation and pupil dilation indicate goal-based predictions whereas anticipatory looking indicated regularity-based predictions?

One possibility considered by Daum et al. (2012) is that anticipatory looking requires rapid computation of the goal and its consequences for movement. It may be that 9-month-olds simply cannot compute the goal in the few hundred milliseconds available for anticipatory looking. This would fit with Daum et al.'s finding that between the ages of 9 months and 3 years, infants' anticipatory looking gradually becomes more adult-like in showing increasingly strong evidence of goal-based action predictions. It would also explain why infants in the first year of life rely on statistical information in anticipating the cow's behaviour in Paulus et al.'s (2011) study, which measured anticipatory looking, but respond in line with the Teleological Stance in Gergely et al. (1995) which used habituation: as in Daum et al.'s (2012) study, infants in Paulus et al.'s (2011) study may simply not have had enough time to incorporate information about goals into action predictions.

But things are not so straightforward. Contradicting the view that 9-month-olds simply cannot compute the goal in the few hundred milliseconds available for anticipatory looking, some researchers have argued that 9- and even 6-month-olds can show evidence of goal-based action predictions in anticipatory looking (for example, Kochukhova and Gredebäck 2010; Green et al. 2016; Cannon and Woodward 2012).606 If this is right, methodological considerations cannot, or cannot fully, explain why infants in the first year of life rely on statistical information in anticipating the cow's behaviour in Paulus et al. (2011).

An alternative explanation is needed. It seems that we do not yet have a rich enough theoretical framework to fully make sense of how infants' use of different kinds of information in anticipating actions differs from that of adults.607 The first puzzle about goal tracking remains a puzzle.

10.7 A second puzzle: acting and tracking

Infants' abilities to track the goals of others' actions appear to be related to their own abilities to act, at least in the first 9 months of life.

How do we know? Part of the evidence for this claim comes from studies in which infants' abilities to act are enhanced. Three-month-olds cannot typically grasp objects or manipulate them with their hands. Needham, Barrett and Peterman (2002) put 'sticky mittens' on 3-month-old infants and let them play with blocks and other toys. By wearing the mittens, these infants could pick up and manipulate the toys in ways normally impossible for 3-month-olds. The next step was to ask whether extending infants' capacities to act in this way might also enhance their abilities to track the goals of observed actions (Sommerville, Woodward and Needham 2005). To measure goal-tracking abilities, these researchers used Woodward's test with the hand reaching for and grasping a teddy or a ball (see Section 10.2 and Figure 10.3). One group of infants first played with objects while wearing the 'sticky mittens' and then took part in the goal tracking test, whereas another group were tested on goal tracking before getting to play while wearing the mittens. Sommerville, Woodward and Needham (2005) reasoned that since playing while wearing the mittens enhances action abilities, infants who played first should be better at goal tracking. And this is just what they found. Three-month-olds who had not yet played while wearing the 'sticky mittens' gave no sign that they could track the goals of observed actions, whereas those who played first did provide evidence of goal tracking.

Even at 3 months of age, some or all of infants' abilities to track the goals of actions they observe are linked to their abilities to perform actions.

Or are they? A potential objection to Sommerville, Woodward and Needham's (2005) study is that the infants who played while wearing 'sticky mittens' first had spent longer observing actions by the time they took part in the goal tracking experiment than the infants who did the goal tracking experiment first. It is conceivable that the difference in their goal tracking is not due to differences in the infants' abilities to act but to differences in the actions they had recently observed. To address this issue, Sommerville, Hildebrand and Crane (2008) conducted a further study. In this study, 10-month-olds were introduced to a novel tool. Infants in one group were given chance to use the tool themselves, thereby perhaps extending their

action possibilities (compare Costantini et al. 2011). Meanwhile, infants in another group were merely allowed to observe the tool being used without holding it themselves. Both groups of infants were then given a goal tracking test similar to Woodward's test (see Section 10.2 and Figure 10.3) but in which objects were not grasped by hand but with the novel tool. Would the infants track the goals of actions involving the tool? As the Motor Theory of Goal Tracking predicts, only infants who had learnt to use the tool themselves did so.

Further support for a relation between infants' goal-tracking abilities and their abilities to act is provided by studies of anticipatory looking in infancy. When observing a hand that is approaching some objects and about to grasp one of them, infants will, like adults, often look to the target of the action in advance on the hand arriving there (Falck-Ytter, Gredeback and Hofsten 2006). This proactive gaze demonstrates rapid goal tracking. Critically, though, infants only do this when they themselves can perform reaching actions. The eyes of infants not yet able to reach do not arrive on an object to be grasped in advance of the hand grasping it (Kanakogi and Itakura 2011).

Even more strikingly, consider what happens when a hand is approaching two objects of different sizes, as in Figure 11.1 on p. 000. Because a reaching hand will be shaped differently depending on whether a large or a small object is about to be grasped, it is possible in principle to detect which object will be grasped. We know that adults can do this because they proactively gaze to the target of action in advance of the hand reaching it (see Section 11.2). What about infants? Their abilities to grasp objects develops slowly. At some point they start to grasp with the whole hand, and then gradually learn to grasp with fewer and fewer fingers. Ambrosini et al. (2013) studied infants with different levels of ability to grasp. They showed that infants' proactive gaze closely matched their grasping abilities. How well infants could grasp with a whole hand was related to how far in advance (or not) their eyes moved to the target of an observed whole-hand grasping action. And infants' abilities to grasp objects precisely were likewise related to their proactive gazes to the targets of observed actions involving a precision grip.

Overall, much evidence links infants' abilities to perform certain actions to their abilities to track the goals of those actions (see Gredebäck and Falck-Ytter 2015, 593–5, for a review). Of course, the two do not correspond perfectly. There are cases where researchers could find no link.608 Further, there is evidence that 3-month-olds can track the goals of reaching actions although they are not capable of reaching (but only of pre-reaching), and that they can do so even for actions like reaching over a high barrier which they could not do at all (Skerry, Carey and Spelke 2013). And 6-month-olds can track the goals of actions (specifically, phonetic gestures) they are wholly unable to perform (Bruderer et al. 2015). So we cannot say that infants' abilities to track the goals of actions are limited by their abilities to perform corresponding actions: and yet there is clearly some link between abilities to perform actions and abilities to track the goals of those actions.

This makes things more puzzling, not less. Why should there be any relation at all between infants' ability to perform an action directed to a type of goal and their ability to track goals of that type?

An initially tempting idea is that performing an action provides new knowledge of means-ends relations, which in turn enhances goal-tracking abilities (compare Skerry, Carey and

Spelke 2013, 18732). Despite promising a straightforward explanation of the puzzle, there are three obstacles to accepting this idea. We would also need to explain why infants' goal tracking is sometimes but not always linked to their abilities to act. Second, we would need to explain why being able to act rather than merely observing is ever necessary for insight into means-ends relations. And, third, there is some—admittedly limited—evidence that even momentarily preventing infants from acting can impair their goal-tracking abilities (Bruderer et al. 2015). This hints that merely acquiring an ability to act may not be enough for goal tracking: what matters may be your abilities at the moment you are tracking the goal.

There is a further problem for the idea that knowledge of means-ends relations might explain why there should be a relation between infants' abilities to act and to track goals. Many researchers have contrasted scenarios involving genuine bodily actions such as a hand grasping an object (or films of these) with scenarios involving non-bodily movements such as a mechanical claw seizing an object (for example, Woodward 1998; Kanakogi and Itakura 2011). The means-end relations involved are essentially the same across the two kinds of scenario. And yet the researchers generally find that behaviours indicative of goal tracking occur only for the bodily actions and not for obviously non-bodily movements.609

We are left with a puzzle. Exactly how is infants' goal tracking linked to their abilities to act, and why should there be any such link?

10.8 Conclusion

Humans can track the goals of actions from around 3 months of age or earlier (as we saw in Section 10.2 and Figure 10.3, which is roughly when they first manifest abilities to track the behaviours of physical objects (see Chapters 2 and 3). How do infants do this?

In thinking about this question, there are good reasons to focus on pure goal tracking, that is, on goal tracking which does not involve ascribing intentions or any other mental states (see Section 10.3). This is not because we know infants cannot track mental states; in fact, there may be reason to guess that infants at this age can do so (as we will see in Chapter 12). It is rather because pure goal tracking could serve as a foundation, in infants and adults alike, for mental state tracking and social interaction.

The Teleological Stance provides an account of how pure goal tracking is possible in theory. Just as rough generalizations about objects can in theory enable a thinker to segment them, represent their persistence and track their causal interactions (see Section 2.3), so also rough generalizations relating actions to goals can in theory enable a thinker to track the goals of action (as we saw in Section 10.4). But is the Teleological Stance descriptively adequate? Does it accurately describe infants' goal-tracking abilities? In attempting to answer this question, we run into two puzzles.

The first puzzle arises because infants' responses are sometimes in line with the Teleological Stance and sometimes appear to be based on statistical regularities. By itself, this is not puzzling: there are good reasons to make use of statistical regularities in goal tracking (see Section 10.5). But in order to have a refutable theory, we need to understand something about when and why infants respond in line with one or another approach. What factors cause

infants in the first 9 months of life to prioritize statistical regularities, and when will they respond in line with the Teleological Stance?

The second puzzle arises because infants' goal-tracking abilities appear to bear a complex and so far unexplained relation to their abilities to act. Sometimes, but not always, enhancing or impairing infants' abilities to perform an action correspondingly enhances or impairs their ability to track the goals of those actions (as we saw in Section 10.7). Before we can conclude that the Teleological Stance is descriptively adequate, we should understand how abilities to track goals are linked to abilities to act and why there should be any such link.

In short, we seek a theory of goal tracking in the first 9 months of life that will enable us to resolve both puzzles. This will eventually provide a basis for understanding the role of goal tracking in the developmental emergence of knowledge.

Notes

11 A theory of goal tracking

As understanding actions is fundamental for profoundly social animals, it is should come as no surprise that goal tracking appears early in development, perhaps from around 3 months of age. But when we attempt even so much as to describe infant goal tracking, we immediately encounter two puzzles (see Chapter 10). We need to understand why infants sometimes rely on statistical regularities and sometimes on the Teleological Stance; and we need to understand how, and why, their goal tracking is linked to their abilities to act. To this end, we need a theory of goal tracking.

As in the case of tracking physical objects and their interactions (see Part I), it makes sense to pick an uncomplicated theory as our starting point and only complicate matters as needed.

11.1 The Simple View

If infants use the Teleological Stance, do they thereby come to know facts about the goals of actions? Csibra and Gergely appear to hold that they do. In their view, applying the Teleological Stance is a matter of reasoning explicitly about actions, about the constraints under which they are performed and about their goals. They stress continuities between goal tracking in infants and explicit reasoning in adults (Csibra and Gergely 1998; Gergely and Csibra 2003), and they describe applying the Teleological Stance as a matter of using knowledge in drawing inferences (Csibra and Gergely 2013). Woodward (1998, 31) also holds that infants' goal-tracking abilities involve knowledge of action.

Let us recycle a label and call Csibra and Gergely's view The Simple View:

The principles comprising the Teleological Stance are things we know or believe, and we are able to track goals by making inferences from these principles.

This view is theoretically coherent. Recall Figure 10.4. You or I can use the principles comprising the Teleological Stance to reason explicitly about the goal of the movements represented in Figure 10.4. According to Csibra and Gergely, what we are doing in this case is what infants (and adults) are doing whenever they apply the Teleological Stance. It follows that infants who apply the Teleological Stance thereby come to know facts about the goals of actions.

One consequence of this view is that we cannot appeal to abilities to communicate with language, nor to rich forms of social interaction in explaining how humans first come to know simple facts about the goals of actions. Instead such an explanation would have to draw on experiences available in the first three months of life, and on innate capacities.

But is the Simple View correct? Unlike in the case of physical objects, I know of no case in which the Simple View concerning the goals of actions generates incorrect predictions. (Perhaps because there is less research.) There is, however, an alternative to the Simple View and one that is arguably better supported. As in the case of physical objects, identifying a compelling alternative to the Simple View is best done by considering how adults track the goals of actions.

11.2 The Motor Theory of Goal Tracking

Adults' eyes show that they are capable of tracking the goals of others' actions extremely rapidly, in mere fractions of second. How do we know? When an adult reaches for an object, her eyes are not typically on her hand but on the object she is reaching for. The same is true of someone observing an action: the observer's eyes are typically on the object someone is reaching for, rather than on the reaching hand (Flanagan and Johansson 2003). These eye movements are proactive: they occur in advance of the action and reveal anticipation of how the action will unfold in the observer. We know that these proactive eye movements reflect goal tracking because they can be influenced by information about the kind of action being observed. For consider that whether someone will grasp an object with her whole hand or between her finger and thumb can often be seen in the shape of the hand, even while the hand is still some distance from the object (see Figure 11.1). When observing someone reaching to grasp one of two different-sized objects, adults will proactively look to the object indicated by the grasping hand's shape (Ambrosini, Costantini and Sinigaglia 2011). This proactive look indicates goal tracking.

How are adults able to track the goals of actions they observe so rapidly? Part of the answer involves motor representations. These are the representations characteristically involved in preparing, performing and monitoring sequences of small actions such as grasping, transporting and placing a fragile egg. We know that performing such actions involves representations of some kind because facts about how subsequent parts of the action will eventually unfold influence how earlier parts of the action are performed (for example, Kawato 1999; Zhang and Rosenbaum 2007). We also know that these representations are not any familiar kind of representation such as intentions or knowledge states because doing things like grasping, transporting and placing an egg involves satisfying constraints not

Figure 11.1 Which of the two balls is this person about to grasp? Schematic representation of one of the grasping events used by Ambrosini et al. (2013).

Figure 11.1

normally considered in explicit practical reasoning. This is the reason for postulating another kind of representation, the motor representation.

Although they were postulated to explain facts about the preparation and performance of actions, it turns out that motor representations live a double life. Motor representations concerning a particular type of action are involved not only in performing an action of that type but also sometimes in observing one. That is, if you were to observe someone grasp one of the two objects in Figure 11.1, motor representations would occur in you much like those that would also occur in you if it were you—not her—who was doing the grasping (Rizzolatti and Sinigaglia 2008, Sinigaglia 2016).

Why do motor representations live a double life? Why might they occur in someone who is not acting other than in observing an action? Consider again someone observing another reaching for a ball in order to grasp it, as in Figure 11.1. Suppose you were to interfere with her ability to represent actions involving the hands motorically, either by tying her hands (as Ambrosini, Sinigaglia and Costantini 2012 did) or by using transcranial magnetic stimulation to temporarily suppress neural activity in bits of the brain most linked to motor representation (as Costantini et al. 2014 did). In both cases you would drastically reduce or even eliminate the proactive gaze. This suggests that motor representations concerning the observed action can facilitate rapid goal tracking in adults. Which is puzzling. Motor representations were postulated to explain action performance. How could such representations also facilitate goal tracking?

An answer is given by what I shall call the Motor Theory of Goal Tracking. In short, this theory says that some pure goal tracking is acting in reverse (see Figure 11.2). More carefully, the

Figure 11.2 The Motor Theory of Goal Tracking

Source: Sinigaglia and Butterfill (2016), Figure 1.

idea is this. In observing an action, all kinds of outcomes may be represented motorically in you more or less simultaneously. These motor representations trigger processes associated with preparing for, performing and monitoring actions, just as they would if it were you, not her, who was acting. As when you are actually acting, these processes lead to expectations concerning which movements you will observe and the sensory effects of these movements. The extent that these predictions are incorrect determines the probability that the motor representation triggering them will be dropped. In many (but not all) cases, this will ensure that motor representations of outcomes other than the goals of the observed action are less likely to be sustained in you than motor representations of the action's goal. And so it is that motor processes in you, as the observer of an action, can ensure, in a limited but useful range of circumstances, that outcomes represented motorically in you are goals of the action you are observing (Sinigaglia and Butterfill 2016).610

The Motor Theory of Goal Tracking seems prone to generating confusion. To avoid some common causes of confusion, consider three points. First, the Motor Theory does not postulate that motor processes operate ‘in reverse’: it relies only on the idea that motor processes take representations of outcomes and compute behavioural and sensory outcomes (as can be seen from the directions of the arrows in Figure 11.2). Second, the Motor Theory does not depend on first identifying an outcome to which an observed action might be directed. This is unnecessary because multiple means–ends computations can occur simultaneously, or at least rapidly enough for action preparation to involve selection on the basis of multiple means-ends computations (for example, Wolpert, Miall and Kawato 1998). Consequently goal tracking that involves motor processes can begin with a wide range of candidate outcomes.

Third, the Motor Theory does not entail that all goal tracking involves motor processes. It is consistent with the fact that some goal tracking can be achieved in other ways, for example, through deliberation. The claim is only that motor processes and representations can enable pure goal tracking.

11.3 The Motor Theory and the Teleological Stance

How is the Motor Theory of Goal Tracking related to the Teleological Stance?

The Teleological Stance provides the basis of a formally adequate account of goal tracking in infants and adults: disregarding limits on memory, attention or processing speed and the like, someone who took the principles comprising the Teleological Stance to be true and made appropriate inferences from these together with observations of bodily configurations and joint displacements could reliably—although not invariably, of course—reach correct conclusions about the goals of actions (see Section 10.4).

The Motor Theory depends on the assumption that the Teleological Stance is indeed formally adequate: it provides an account of how motor representations and processes could, within limits, implement the computations described by the Teleological Stance.

The Motor Theory also provides a possible link between the Teleological Stance and the mind of an individual. Suppose we ask the question: What links the mind of an individual to the principles comprising the Teleological Stance? One such account of the link is given by the Simple View: the principles comprising the Teleological Stance are known by the adults and they track goals by making inferences from these principles. Another possible account is suggested by the Motor Theory of Goal Tracking: the principles comprising the Teleological Stance characterize how motor processes in that individual enable goal tracking.611

Can we use the Motor Theory to characterize infants' earliest ability to track the goals of actions? Consider what I will call the Developmental Motor Conjecture:

In the first nine months of life, all pure goal tracking is explained by the Motor Theory of Goal Tracking. Other goal-tracking processes emerge later in development.612

This conjecture has a striking virtue. As we saw, in the first nine months of life, infants' goal tracking abilities bear an interesting, hard-to-pin-down relation to their abilities to act (see Section 10.7). The Developmental Motor Conjecture implies that these infants' goal-tracking abilities should be limited by their abilities to represent actions motorically. Since abilities to represent actions motorically are loosely related to abilities to perform those actions, the Developmental Motor Conjecture has the potential to explain why infant goal-tracking abilities should be so limited.

There is just one problem. On the face of it, the Developmental Motor Conjecture appears to be completely untenable. We know that, as a general rule, movements of simple geometric shapes such as those used in Gergely et al.'s (1995) experiment (see Section 10.2) are unlikely to be represented motorically. And one implication of the conjecture is that infants'

goal-tracking abilities should be limited by their abilities to represent events motorically. And yet, as we saw, there seems to be abundant evidence that infants can track the goals of actions performed by geometric shapes, cartoon fish, and the like.

This appears to be a compelling reason to reject the Developmental Motor Conjecture.613 But appearances can be deceptive. Despite appearing obviously wrong, the Developmental Motor Conjecture may actually be correct. To see why, we first need to distinguish between targets and goals.

11.4 Target vs goal

The orthodox view we have been uncritically following so far ignores a critical distinction between goals and targets. The target or targets of an action (if any) are the things towards which it is directed. If the goal of an action is to kick a particular football, this football is the action's target. To specify a target of an action is to partially specify one of its goals. But more is required to fully specify a goal, of course. A goal typically involves a type of action—kicking rather than smashing, say. It may also involve one or more manners of action—discreetly, firmly, and precisely, for example—and perhaps more besides (see Figure 11.3).

Why does the distinction between targets and goals matter? In adults, there is a capacity to track targets only. This is called perceptual animacy, the detection by broadly perceptual processes of animate objects and their targets. To illustrate, consider an experiment by Gao, Newman and Scholl (2009 experiment 1).614 Adults were shown a display which contained some moving circles. In some cases the circles moved independently of each other, but in other cases there was a ‘wolf’ which chased a ‘sheep’ with varying degrees of subtlety (see Figure 11.4). The adults' task was simply to detect the presence of a wolf. Gao, Newman and Scholl (2009) established that adults can do this providing the chasing is not too subtle. In further experiments, they also showed that adults' abilities to perceptually detect chasing depend on several cues including whether the chaser ‘faces’ its target (‘directionality’) and how directly the chaser approaches its target (‘subtlety’).

Perceptual animacy is commonly interpreted as a case of goal tracking.615 But it is important to make a distinction. Perceptual animacy involves tracking targets only. The type and manner of action are unspecified and irrelevant, as are any further features of outcomes (see Figure 11.3). So perceptual animacy only counts as goal tracking in an attenuated sense.

Figure 11.3 Fully specifying a goal can involve giving a type of action, a target, some manners of action, and more

Figure 11.3

typetarget(s)manner(s)relation(s)...
push / grasp / kick / ...the ball / the tomato / the frog / ...softly / discreetly / energetically / ...before Ayesha / while it is red / in my turn / ......

Figure 11.4 Schematic representation of a chasing event in an experiment on perceptual animacy

Source: Gao, Newman and Scholl (2009), Figure 2.

The detection of animacy appears to be a broadly perceptual phenomena since it depends on areas of the brain associated with vision and influences how perceptual attention is allocated (Scholl and Gao 2013), irrespective of your beliefs and intentions (van Buren, Uddenberg and Scholl 2016). Perceptual animacy also appears to depend on simple cues and heuristics involving motion trajectories like directionality and subtlety, as we say. Applying these heuristics would be detrimental to proper goal tracking. It is therefore unlikely that perceptual animacy is correctly described by the Teleological Stance. Whatever broadly perceptual abilities underpin perceptual animacy are probably distinct from those which enable goal tracking proper.

To avoid confusion, let us distinguish merely tracking targets from proper goal tracking, which is goal tracking that is not merely target tracking. Perceptual animacy involves merely tracking targets.

Are infants in the first nine months of life capable of proper goal tracking? Nearly all of the studies we considered in Chapter 10 do not distinguish whether infants in the first 9 months of life are merely tracking targets of actions or whether they are properly tracking goals. They do not show, for instance, that infants can identify the type of an action—whether it is a grasping or a pushing action, say. To say that infants can properly track goals and not merely targets implies, minimally, that they can distinguish both the target and the type of an action.

So can infants also distinguish between two actions which are directed to the same target but differ in type? To answer this question, we would ideally have pairs of scenarios in which the target of an action is kept constant while the type of action varies. To the extent that subjects respond appropriately to the difference in type of action, we can be confident that they can distinguish actions not just by their targets but also by their types.

Behne et al. (2005) created just such pairs of contrasting scenarios (albeit for a different purpose). In one of their contrasts, an experimenter holds a ball out for an infant to grasp and then either ‘accidentally’ drops it or teasingly pulls it back. So in each case there is a

goal-directed action involving the ball, but in one case the goal of the action is to pass the ball to the infant whereas in the other case the goal is to tease the infant. Behne et al. (2005, Study 2) found that 9-month-olds (but not 6-month-olds) consistently and appropriately discriminated between these scenarios by, for example, banging more when the ball was ‘accidentally’ dropped than when it was teasingly retracted. This and other research (for example, Ambrosini et al. 2013, discussed in Section 10.7; see also Kochukhova and Gredebäck 2010; Green et al. 2016) suggest that, at least from 9 months of age, infants can indeed distinguish both the type and target of a goal-directed action.

The fact that infants in the first nine months of life are capable of proper goal tracking indicates that their abilities cannot be entirely a consequence of perceptual animacy. But this leaves open the possibility that some of the experimental observations standardly considered to support goal tracking in infancy may in fact be explained by the perceptual detection of animacy. Reflection on this possibility enables us to develop a theory and solve the twin puzzles about infants' goal tracking.

11.5 A dual process theory of goal tracking

In the first year of life, infants' abilities concerning physical objects appear to involve at least two distinct kinds of process, one broadly perceptual and the other broadly motoric (see Chapter 6). Perhaps something similar is true of their abilities concerning actions.

The Developmental Motor Conjecture states that all goal tracking in the first nine months of life is explained by the Motor Theory of Goal Tracking. As we saw (in Section 11.3), this Conjecture has the potential to explain how and why infants' goal-tracking abilities are linked to their abilities to perform actions. However, it faces an objection: 9-month-olds seem to exhibit goal tracking when confronted with scenarios involving self-propelled balls (Csibra and Gergely 1998) or cartoon fish (Daum et al. 2012), which are unlikely to trigger motor processes. But what if the supposed goal tracking in these situations is mere target tracking, and what if it were a consequence of perceptual animacy?

Recall Gergely et al.'s (1995) ground-breaking study with the balls (see Section 10.2 and Figures 10.1 and 10.2). In principle, this effect is no less likely to be a consequence of perceptual animacy than of goal tracking. There are hints that infants in the first year of life can perceptually detect animacy (Rochat, Striano and Morgan 2004). And because Gergely et al.'s (1995) effect appears to depend on cues to animacy (Schlottmann and Ray 2010), an interpretation in terms of perceptual animacy cannot currently be ruled out.

These reflections motivate considering what I will call Conjecture MP:

In the first nine months of life, all proper pure goal tracking is explained by the Motor Theory. Other pure goal-tracking processes emerge later in development. Further, a mere target-tracking process is also present in these infants. This process is identical to perceptual animacy in adults. And appearances that these infants' pure goal-tracking abilities are not limited by what they can represent motorically are misleading: they are due to mistaking mere target tracking for proper goal tracking.616

This conjecture provides a dual process theory of goal tracking. It invokes distinct kinds of motor and perceptual processes, which are involved in distinct kinds of tracking (namely, goal tracking proper and mere target tracking).

The objection to the Developmental Motor Conjecture was that 9-month-olds can track the goals of animate balls and cartoon figures, whereas that conjecture implies they could not. Conjecture MP overcomes this objection by allowing that infants' goal tracking involves two distinct kinds of process, one perceptual and the other motoric.

Although Conjecture MP is yet to be tested directly, it does generate readily testable predictions. One source of predictions is the fact that perceptual animacy involves only broadly perceptual processes whereas goal tracking is facilitated by motor processes. Further, perceptual animacy and goal tracking have distinct signature limits. Goal tracking can be impaired by tying an observer's hands or otherwise interfering with her capacities to represent actions motorically (see Section 11.2). By contrast, the perceptual detection of animacy can be impaired by manipulating unrelated factors like directionality and subtlety (Gao, Newman and Scholl 2009). A final source of predictions is the distinctive effects of perceptual animacy on attention (van Buren, Uddenberg and Scholl 2016).

Conjecture MP is quite likely to be wrong, just like any other bold and as yet untested conjecture is. But it makes a certain kind of sense. Given the importance of action, we should expect that even in the first few months of life infants will manifest a variety of abilities in observing actions and that these will depend on a mix of perceptual and motor processes which are identical, or closely related, to those found in adult humans.

11.6 Puzzles solved?

In Chapter 10 we encountered two puzzles about goal tracking in the first nine months of life. Conjecture MP offers a possible explanation of both puzzles.

The first puzzle was why infants sometimes respond on the basis of statistical regularities whereas at other times their responses are in line with the Teleological Stance. Conjecture MP suggests a candidate answer, as summarized in Table 11.1. This candidate answer depends on three further, plausible assumptions. The first is that where scenarios involve human agents, either perceptual animacy processes will not occur or else they will tend to be dominated by motor processes in infants' responses. The second assumption is that, in infants at least, perceptual animacy processes can drive looking times and pupil dilation but

Table 11.1 Which processes will tend to dominate 9-month-olds' responses to a scenario involving actions? According to Conjecture MP, the answer depends on both scenario type (rows) and response type (columns)

Habituation, violation-of-expectation and pupil dilationAnticipatory looking
Observing infant-possible actionsMotor representationsMotor representations
Observing infant-impossible actionsPerceptual animacyStatistical regularities

not anticipatory looking. The third assumption is that infants' responses will typically only be dominated by statistical regularities when no form of goal tracking is possible. Given these further assumptions, Conjecture MP implies that infants' responses to action scenarios should manifest reliance on statistical regularities when they are observing actions they cannot perform and responding in a way that manifests anticipation of how the action will unfold. And, as we saw in Section 10.5, this is just what Daum et al. (2012) and others have observed.

The second puzzle about action was to understand how infants' goal tracking is linked to their abilities to act and why there should be any such link (see Section 10.7). According to Conjecture MP, the kind of link involved depends on whether the infant is merely tracking targets or engaged in proper goal tracking. Because proper goal tracking in the first nine months of life depends exclusively on motor processes, infants' goal tracking should be limited by their abilities to represent observed actions motorically, much as motor-based goal tracking is limited in adults (see Section 11.2). But because mere target tracking depend on perceptual animacy rather than any kind of motor process, mere target tracking should not be limited in the same way. Given the assumption that scenarios involving self-propelled balls (Csibra and Gergely 1998) or cartoon fish (Daum et al. 2012) involve mere target tracking, it is no surprise that infants' goal-tracking abilities in these scenarios are not constrained by what they can represent motorically.

We encountered these two puzzles in asking whether the Teleological Stance is descriptively or explanatorily adequate. Just here Conjecture MP has a subtle consequence. If this conjecture is correct, when infants are merely tracking targets, the Teleological Stance should be neither. By contrast, when they are engaged in proper goal tracking, it should be both descriptively and explanatorily adequate.

11.7 Conclusion

The Teleological Stance demonstrates the theoretical possibility of pure goal tracking (Section 10.4). Humans are capable of pure goal tracking from around 3 months of age or earlier and they manifest extremely rapid pure goal tracking in anticipatory looking from around 6 months of age (see Sections 10.2 and 10.7).

Do these early goal-tracking abilities involve coming to know simple facts about the goals of particular actions? According to the Simple View, the principles comprising the Teleological Stance are things we know or believe, and we are able to track goals by making inferences from these principles. So tracking goals involves whatever kind of inference is involved when a detective figures out who committed a murder, and it does result in knowledge about the goals of particular actions.

The Simple View correctly characterizes some goal tracking in human adults (see Section 11.1), so it is not theoretically incoherent to suppose that it might also characterize younger humans' goal tracking too. But there is more to the story about goal tracking in adults. There is evidence for the view that goal tracking in adults is sometimes a consequence of motor

processes and representations only (see Section 11.2). We should therefore at least consider the possibility that motor processes underpin some infant goal tracking.

A first, too simple attempt to articulate a view along these lines is the Developmental Motor Conjecture. This conjecture states that all pure goal tracking in the first nine months of life is explained by the Motor Theory of Goal Tracking. As it stands, this conjecture is untenable because it implies, incorrectly, that infants cannot track goals involving animate balls and cartoon fish (see Section 11.3).

To make progress in understanding goal tracking, it may be essential to distinguish goals from targets, proper goal tracking from mere target tracking, and motor processes from perceptual animacy (see Section 11.4). With these distinctions, we can construct a conjecture that provides natural solutions to two puzzles about action as well as generating some readily testable predictions (see Section 11.6). This is Conjecture MP, which holds that goal tracking in the first nine months of life involves a combination of perceptual animacy and motor processes.

Whether we stick with the Simple View or accept a conjecture along these lines has implications for how the emergence in development of knowledge of action might be explained. The Simple View entails that infants who can track goals know simple facts about the goals of particular actions. If we accepted the Simple View, we would therefore have to conclude that the emergence of such knowledge depends only on experiences gathered in the first three months of life and on any innate capacities. And in invoking goal-tracking abilities to explain feats like communication later in infancy, we would be appealing to knowledge states. By contrast, goal tracking that is correctly characterized by Conjecture MP does not necessarily result in knowledge concerning the goals of particular actions. According to this conjecture, goal tracking in the first nine months of life involves not knowledge states but perceptual and motor representations, and these are not inferentially integrated with knowledge. We may therefore suppose that knowledge of the goals of particular actions emerges much later in development, and may depend on rich forms of social interaction and perhaps abilities to communicate with words.

One consequence is that goal tracking can serve as a building block in explanations of how humans develop abilities to engage in rich forms of social interaction (see Chapter 15), and track others' mental states (see Chapters 12 and 14). Philosophers tend to focus on these capacities and ignore the pure goal tracking that underpins them. This is probably a mistake. In tracking the goals of actions, you are taking the step from mere joint displacements and bodily movements to representations of the outcomes around which these are organized (see Section 10.3). This is arguably the hardest and most significant step in making sense of others, their thoughts and the things they say. It is like the step from specifying the states of particular electronic circuits in a computer to being able to use machine language, which abstracts from physical considerations like voltage while closely reflecting the architecture of the underlying hardware. There is much further abstraction to come, of course. But the transition from electronics to machine language, or from kinematics to goals, is a foundation for everything else.

Notes

12 Mind

The puzzle

How do humans first come to know facts about others' mental states? How, for instance, do they come to know that Ayesha believes, falsely, that she and Beatrice will still be able to catch a bus home even if they delay leaving the party? By the end of this chapter you should understand the facts that make answering this question difficult; you should have a sense of how far away we all are from being able to answer it definitively; and you should be familiar with some attempts to answer it. You should also be able to evaluate novel answers to the question about how humans first come to know facts about others' mental states.

The focus of this chapter is what is often called mindreading or using a theory of mind. Mindreading is the process of identifying a mental state as a mental state that some particular individual, another or yourself, has. To say someone has a theory of mind is another way of saying that she is capable of mindreading.701 As is clear from these definitions, having a theory of mind does not necessarily involve having a theory. Similarly, talking about mindreading does not commit one to the idea that mental states can literally be read, nor to magical thinking about knowledge of others' minds. Despite their inappropriate connotations, we are stuck with the terms ‘theory of mind’ and ‘mindreading’ because they are so widely used.

Most research on mindreading concerns belief rather than, say, desire or intention. This is partly because it was quite widely held that the mark of a mindreader was an ability to represent false beliefs (see Perner 1991). Without being committed to this claim, I shall also focus on belief and belief-like states in this chapter. This is a practical matter: even if you don't think there is anything special about understanding belief as opposed to understanding other mental states, understanding the research on belief is necessary for thinking about how humans come to know facts about mental states generally.

12.1 All about Maxi

Wimmer and Perner (1983) set out to determine when humans can know facts about others' beliefs. They told children a story like this: ‘Maxi puts his chocolate in the BLUE box and leaves the room to play. While he is away (and cannot see), his mother moves the chocolate from the BLUE box to the GREEN box. Later Maxi returns. He wants his chocolate.’ They then asked the children, ‘Where will Maxi look for his chocolate?’

The results are amazing (see Figure 12.1). At least, they are amazing to most adults who don't already know the results of this experiment and have at least a vague sense of what 4- and 5-year-olds are like. They tend to find it amazing that children under around 3 or 4 years of age systematically give the wrong answer, and that children do not typically reach adult levels of performance until at least 4 or 5 years of age.

Given how amazing the results are, you might be tempted to think that something is wrong with the experiment. Could it be that children are confused by the story or the question, or find it tricky for reasons that have nothing to do with their understanding of mental states? It turns out that children's performance is consistent across many variations of the false belief task. Instead of asking them to predict Maxi's action, you can ask them to predict his desire (Astington and Olson 1991), or you can show them how he acted and ask them to retrodict his belief (for example, Wimmer and Mayringer 1998). Or you can use tasks which actively involve children, testing their propensities to perform deceptive actions (for example, Chandler, Fritz and Hala 1989) or to lie (for example, Polak and Harris 1999; Talwar, Gordon and Lee 2007). You can give children non-verbal false belief tasks which involve choosing something based on another's false belief (for example, Call and Tomasello 1999; Krachun et al. 2010) or placing a bet (Ruffman et al. 2001). You can even ask children about their own (past) false beliefs rather than somebody else's (Gopnik and Astington 1988). Whichever of

Figure 12.1 The results of Wimmer and Perner's (1983) false belief task

Source: Drawn from Wimmer and Perner (1983).

these ways you test children's understanding of false belief, you will get basically the same results Wimmer and Perner got.

In fact, Wimmer and Perner's finding is one of the most replicated in developmental psychology. A meta-analysis of 178 studies confirms their basic finding: sometime around their fourth birthday, children go from systematically acting and talking as if false belief were an impossibility to being able to recognize false beliefs (Wellman, Cross and Watson 2001).

This is not to say that every experiment gives exactly the same results. If you look at the results of different studies, you can see that there is plenty of variation. Consider what happens when children around 40 months of age are tested, for example. Some studies find that nearly all children of this age give incorrect answers to a false belief task, whereas other studies find that nearly all children of this age give correct answers; and most studies lie somewhere between these extremes. If you looked at one or two studies in isolation, you might think that they provide contradictory findings. But in fact there is no contradiction. The different studies used different versions of the false belief task. And different versions include different combinations of features which made the task easier or harder. For example, changing the question from ‘Where will Maxi look for his chocolate?’ to ‘Where will Maxi look first for his chocolate?’ can improve performance (Siegal and Beattie 1991). Similarly, changing the task so that it involves an element of deception, or so that children are more involved because they interact with Maxi (or the other protagonist), makes the task a bit easier too. But these changes make the task easier for children of all ages. They do not change Wimmer and Perner's basic finding that there is a transition from systematically incorrect to systematically correct performance in typically developing children.702

The ages at which typically developing children first pass false belief tasks can vary quite a bit. There are differences between children in different countries, as Figure 12.2 shows. Several components of a child's linguistic and communicative abilities also appear to affect when she will first pass a false belief task (Milligan, Astington and Dack 2007; Kovács 2009). Also children who are better at inhibiting natural tendencies, switching tasks and holding things in working memory tend to pass false belief tasks at a younger age than children who are otherwise similar (Devine and Hughes 2014). Another predictor of when children first pass false belief tasks is the quality of their social interactions (Hughes et al. 2006). Children with mothers who more frequently give explanations involving mental states or more frequently talk about minds tend to pass false belief tasks at a younger age (Peterson and Slaughter 2003; Moeller and Schick 2006), and children with a variety of siblings (but not twins) also tend to pass at a younger age than otherwise similar children (Ruffman et al. 1998; Cassidy et al. 2005). It is even possible to get children to pass false belief tasks earlier than they would otherwise by giving them training relevant to understanding and reasoning about mental states (for example, Slaughter and Gopnik 1996; Lohmann and Tomasello 2003; Wellman and Peterson 2013).

At this point we appear to be in a good position to start figuring out how humans first come to know facts about beliefs and other mental states. The research just mentioned tells us that a variety of cognitive, linguistic and social factors are associated with success on false belief tasks. Deeper exploration might allow us to disentangle correlation from causation, perhaps leading us to the view that humans first come to know facts about mental states in

Figure 12.2 Performance of children in different countries on false belief tasks

Source: H. Wellman, Cross and Watson (2001), Figure 7.

roughly the way that they first come to know facts about what written words mean: they learn about mental states like belief through social interaction. (Heyes and Frith 2014 develop a version of this view.)

But there is another body of evidence which appears to conflict dramatically with this line of thought.

12.2 Infants track false beliefs

Recall Wimmer and Perner's experiment with the story about Maxi. In this experiment children are asked to say where Maxi (who manifestly has a false belief) will look for his chocolate. What happens if, before asking this question, the experimenter simply muses, as if thinking aloud, ‘I wonder where Maxi will look for his chocolate’? When given such a prompt, children (and adults) tend to look to the location where they expect something to happen. Where children look in response to the prompt can therefore reveal what they expect to happen. And when Clements and Perner (1994) measured children's anticipatory looking, they found that 3-year-olds would look to where Maxi should go if he had a false belief. (Actually they used a different scenario with a mouse and some cheese, but it has the same structure as the Maxi

story.) Yet when asked the question about where Maxi will go, these same children would systematically, and confidently, predict that Maxi will act as if he had a true belief (Garnham and Ruffman 2001; Ruffman et al. 2001).

Why are these results exciting and puzzling? It isn't primarily about age—it isn't merely that they show that 3-year-olds can track false beliefs. Rather the exciting feature is that for individual children observing a single scenario, anticipatory looking and verbal prediction give contradictory indications about what they expect to happen next. Their anticipatory looking suggests they expect Maxi to act on a false belief and go one way, whereas their verbal predictions suggest that they expect Maxi to act as if false belief were impossible and so to go the other way. Since Maxi can't go both ways, Clements and Perner's 3-year-olds appear to have contradictory expectations.

The case for accepting that eye movements reveal an ability to anticipate actions based on false beliefs is strengthened by evidence that both 2-year-olds' (Southgate, Senju and Csibra 2007) and also adults' (Low and Watts 2013) anticipatory looking shows a similar pattern. So 3-year-olds' anticipatory looking is not a quirk: individuals of all ages look as if they are anticipating actions based on false beliefs.

And it is not only in their anticipatory looking that children younger than 3 years of age manifest abilities to form expectations concerning actions based on false beliefs. Onishi and Baillargeon (2005) made a discovery that has transformed our understanding of how humans first come to know facts about minds. They created a violation-of-expectation version of the false belief task which enabled them to test 14-month-olds. (The violation-of-expectation method was explained in Section 2.2.) As in Wimmer and Perner's task, children observe a sequence of events in which a protagonist (like Maxi) manifestly acquires a false belief about the location of an object. Of course, the young age of the subjects means that the sequence is acted out for them rather than told as a story. And instead of asking their infants a question, Onishi and Baillargeon had their protagonist continue to the end of the story, either by reaching to the wrong location or by reaching to the correct location. If infants expect the protagonist to act in accordance with her false belief, then they should look longer when the protagonist manifestly has a false belief but nevertheless acts as if she had a true belief. And this is exactly what the infants did. As you can see in Figure 12.3, 14-month-olds show the converse pattern of looking when the protagonist has a true belief.

Following Onishi and Baillargeon's ground-breaking experiment, many have replicated and extended their findings. You can change the scenario so that it involves false belief about the contents of a container rather than about the location of an object (He, Bolz and Baillargeon 2011), or so that it involves relating beliefs to expressions of emotion rather than to instrumental actions (Scott 2017), or so that it involves inferring belief from verbal communication (Scott et al. 2012) or from action (Träuble, Marinović and Pauen 2010). Despite these quite radical changes to the tasks, essentially the same results are found in each case: children in their second or third year of life can form different expectations depending on what another believes. Taking these results together with experiments measuring anticipatory looking, we have evidence that even 1-year-olds can form expectations about actions based on false beliefs.

Individual 2- and 3-year-old children observing a single scenario involving a protagonist like Maxi acting on a false belief appear to have inconsistent expectations concerning what the

Figure 12.3 Fourteen-month-olds expect people to act in line with their beliefs, whether true or false

Source: Drawn from Onishi and Baillargeon (2005).

protagonist will do. Whereas some responses robustly indicate that the child expects the protagonist to act as if false belief were impossible (see Section 12.1), other responses no less robustly indicate that the child expects the protagonist to act in accordance with her false belief (as we have seen in this section). I shall eventually suggest that these apparently contradictory responses give rise to a deep and difficult puzzle about the nature of mindreading. But before I say what the puzzle is, let us first be sure that its foundations are solid. Do 2- and 3-year-olds really make contradictory responses to scenarios involving false beliefs?

12.3 A replication challenge

One challenge concerns whether studies of infant false belief tracking can be replicated, and indeed whether corresponding studies with adults replicate. A surprising number of findings have turned out to be inexplicably hard to replicate, while other findings have been replicated (for example, Kulke et al. 2017; Powell et al. 2017; Crivello and Poulin-Dubois 2017; Dörrenberg, Rakoczy and Liszkowski 2018; Kulke and Rakoczy 2018; Kulke et al. 2018). Even more confusingly, some findings have been both successfully and unsuccessfully replicated (for example, see Kulke and Rakoczy 2018 on Southgate, Senju and Csibra 2007).

Attempts to explain why some findings can be replicated and others not (and why some findings sometimes but not always replicate) are currently quite varied. Some argue that particular failed replications are due to methodological differences (for example, Buttelmann, Baillargeon and Southgate, n.d.; see Poulin-Dubois et al. 2018, for responses). Others suggest that particular measures, most prominently anticipatory looking, may not be reliable indicators of belief tracking at all (Kulke, Johannsen and Rakoczy 2019, 14). But I think most would agree that any systematic explanation for the pattern of successes and failures in replication attempts is some way off.

Despite the uncertainty, two challenges seem particularly important. First, when various tasks are supposed to measure a single ability, we would normally expect to find signs of convergence in performance across the tasks: that is, those and only those subjects who pass one of these tasks will tend to pass other tasks. Kulke et al. (2017, 2) observe that whereas performance on false belief tasks used to test older children is convergent in this sense, there is little evidence of convergence for false belief tasks suitable for infants; and Poulin-Dubois and Yott (2017) find evidence for divergence.703 Second, Wellman (2018, 741) notes that in tasks typically used with older children, measures of belief tracking are predictive of social skills, whereas there is as yet little evidence that performance on belief tracking tasks used with infants predicts social abilities.

Perhaps future research will overcome these challenges by uncovering evidence of convergence and predictive value. Alternatively, it may be that the challenges are based on an incorrect premise. Given that there is individual (and cultural) variability in the ability measured by false beliefs tasks designed for use with older children, it makes sense to expect convergence and predictive value. But perhaps the ability measured by false beliefs tasks suitable for use with infants does not vary much between individuals or across development.704 If this turned out to be right, we would not necessarily expect performance on tasks suitable for infants to show signs of convergence, nor should it have much predictive value concerning social skills. But this is speculation. Until these challenges are resolved, there is surely room for doubt about whether 1-year-olds really can form expectations about actions based on false beliefs.

What follows rests on a guess. My guess is that even 2- and 3-year-olds really can track beliefs. I thought there was already a case for this guess 20 years before I wrote this chapter (Butterfill 2001). And even taking seriously challenges raised by patterns of success and failure in replication studies, on balance, the evidence in favour of this guess has grown since then.

If we accept this guess, we are immediately confronted with a complementary challenge to the idea that 2- and 3-year-olds make contradictory responses to scenarios involving false beliefs. Perhaps the tasks on which 2- and 3-year-olds fail to exhibit belief tracking are just badly designed?

12.4 Methodological defects or truly contradictory responses?

We have just been considering a challenge to the idea that performance tasks suitable for infants really provide evidence for belief tracking in 2- and 3-year-olds. The next, converse challenge concerns whether the tasks on which 2- and 3-year-olds fail to exhibit belief tracking really indicate an absence of belief-tracking abilities.

Consider the responses children make which indicate that they expect Maxi (or whoever) to act as if false belief were impossible. These are typically (but not always) responses to a question directed to the child. Asking children a question invites them to consider multiple possibilities, and it requires them to engage conversationally with an experimenter who is not part of the story. By contrast, anticipatory looking and violation-of-expectation paradigms do

not interrupt children’s engagement with the scenario in these ways. Perhaps, then, children are not really making contradictory responses. Instead it might be that anticipatory looking and violation-of-expectation paradigms provide a sensitive measure of children’s expectations concerning how Maxi will act, whereas the responses which appear to indicate that they expect Maxi (or whoever) to act as if false belief were impossible are merely a reflection of poor experimental design resulting in children becoming distracted.705

But is this right? Can we explain away 2-and 3-year-olds apparently contradictory responses by saying that anticipatory looking and violation-of-expectation paradigms simply use better measures? Two considerations jointly suggest that we cannot. The first consideration is that 3-year-old children succeed on tasks that are very similar to false belief tasks but concern desire, perception or pretence rather than belief (Gopnik and Slaughter 1991; Gopnik, Slaughter and Meltzoff 1994; Custer 1996). In fact, they can even succeed on tasks which require contrasting their own desires with another’s incompatible desires (Rakoczy, Warneken and Tomasello 2007, Study 2). This shows that, by 3 years of age and probably earlier, children can competently make verbal predictions in response to questions, even where doing so involves taking into account facts about others’ mental states. The second consideration is that children between their second and fourth birthdays do not usually answer at random when asked about false beliefs. Instead they systematically and confidently give answers that contradict those that adults usually provide: they answer questions as if there were no such thing as false belief. This consideration shows that children whose anticipatory looking manifests expectations about false beliefs are not all merely confused when they are asked a question.706

Taken together, these two considerations suggest that answers by typically developing 2- and 3-year-olds to questions involving belief are based on a sensible and coherent model of minds and actions in which belief does not feature. This is a model on which how people act is a consequence of what they want (or perhaps of what is desirable for someone in their position), of how the world is, of which parts of it they have engaged with, and perhaps also of their emotions. Someone who wants to eat the chocolate will go to where the chocolate actually is (unless prevented from doing so by never having engaged with the chocolate); and someone who wants to feed a mouse will fetch the cheese. For our purposes, the key feature of these explanations is that they imply 2- and 3-year-olds rely on a model of minds and actions which does not incorporate beliefs. What determines how people act on such a model is not what they believe to be the case but rather what is the case. Several philosophers have argued that a model along these lines is theoretically coherent and enables accurate predictions and explanations in a wide variety of situations (Stout 1996; Gordon 2000). And several developmental psychologists have independently argued that children’s early reasoning about actions is based on such a model.707

On the face of it, then, what prevents most 2- and 3-year-old children from answering questions about false beliefs as adults typically do is that, in answering these questions, they rely on a model of minds and actions in which belief does not feature.

This view is supported by several further considerations. One is that, as already mentioned, such children appear to shift to giving belief-based answers following training that is specifically relevant to understanding and reasoning about mental states (for example, Slaughter

and Gopnik 1996; Lohmann and Tomasello 2003). Another consideration is that children’s abilities to answer questions about false beliefs predict aspects of social competence, and do so independently of factors like their abilities to inhibit natural responses to things and to hold several things in mind (Razza and Blair 2009). These facts about the causes and consequences of typically developing children’s performance on the kinds of false belief task which they usually fail until around 4 years of age are hard to reconcile with the view that their performance is an artefact of poor experimental design. But the facts do make perfect sense if children shift from answering questions as if false belief were impossible to giving answers that take into account false beliefs because they shift from relying on a belief-free model of minds and actions to a model which does incorporate beliefs.

There seems to be little prospect of finding a methodological defect in the false belief tasks that typically developing 2- and 3-year-olds tend to fail.

At this point I suggest we have sufficient evidence to conclude that typically developing 2- and 3-year-olds’ responses to scenarios involving false beliefs really do manifest contradictory expectations. The appearance of contradiction is genuine and not merely an artefact of experimental design.

Note that accepting this modest claim leaves open the much harder question of what underpins these conflicting responses. I have been suggesting that, on the face of it, those responses which indicate an expectation that a protagonist will act in accordance with a false belief involve the use of a model of minds and actions incorporating beliefs. As we will see in Section 13.1, there are challenges to this claim. I have also been suggesting that, on the face of it, those responses which indicate an expectation that a protagonist will act as if false belief were impossible involve the use of a model of minds and actions not incorporating beliefs. There are significant challenges to this claim too—some hold that the responses reflect difficulties in applying a model of minds and actions to certain tasks, and do not reveal features of the model itself (see Sections 13.5 and 13.7). There is, then, plenty of controversy to come. But one thing should not be controversial at all: 2- and 3-year-olds really do tend to make contradictory responses to scenarios involving false beliefs. This leaves us with a puzzle.

12.5 Models

I have been talking about children relying on a model of minds and actions without explaining what this means.

For simplicity, start by thinking about a physical model of a house. This model could be used in various ways. Perhaps initially you use the model to imagine your dream house, but then you win the lottery and use the model as a plan for building an actual house. Later you display the model in the hallway and your house guests use it to find their way around.

A model is something that can serve different purposes. Having a model does not commit you to using it for any particular purpose. (If your dreams are more extravagant than your lottery winnings, you could use a different model when building the actual house.) The model’s usefulness does not depend only on its accuracy: the ease with which it can be used to imagine, build or navigate matters. The best model for a given set of purposes may not be

the most accurate. Further, it can be advantageous to have multiple models of a single thing. For example, building a house can involve creating multiple models. There is nothing wrong with having and using more than one model.

In talking about someone relying on a model of minds and actions, there is clearly no suggestion that they have a physical model. In general, a model is a way some part of aspect of the world could be. So a model of minds and actions is just a way mental aspects of the world could be.

A model is distinct from a theory. A model can be used to make claims about the world, of course (‘this is the bathroom’), but the model itself entails nothing about how the world actually is. By contrast, a theory does (Godfrey-Smith 2005). Models are easily confused with theories because we can use theories to specify models. To describe the model of minds and actions a child or adult relies on, we might give a theory of the mental. Similarly, in stating the Principles of Object Perception we partially characterized a model of the physical that infants (and adults) rely on (see Section 2.3). But of course to say that someone relies on a model is not to say that they have or endorse a theory. The theory is what we as theorists use to characterize the model.708

But why do we need models at all? The answer is easier to see by thinking about the physical case. Suppose we discover that infants in the first four months of life can track briefly occluded physical objects. This discovery raises the question, what are physical objects like from the point of view of the infant? We can answer this question by identifying a model of the physical. The model tells us what the physical world is like from the infant’s point of view (or from the point of view of some process in the infant). Similarly, to say that infants in the first or second year of life can track a belief leaves entirely open the question of how the mental appears to the infant (if it appears to the infant at all). Identifying a model of minds and actions allows us to characterize how minds are from the infants’ point of view (or from the point of view of a process in the infant).

Consider a 2- or 3-year-old who, when asked where Maxi will look for his chocolate (see Section 12.1), confidently tells you that Maxi will look where his chocolate actually is (so not where Maxi believes his chocolate is). Should we therefore conclude that such a child is not relying on a model of minds and actions? Not at all. The child’s answer may be based on a coherent model of minds and actions, one on which facts guide actions. Although adults are often surprised that children answer questions on the basis of facts rather than beliefs, there is nothing obviously wrong with the strategy of using facts to predict and explain actions. Indeed, it is possible to develop a sophisticated and coherent theory of minds and actions along these lines (see, for example, Stout 1996). Children who systematically make predictions about actions as if agents always acted on the relevant facts are plausibly relying on a coherent model of minds and actions which does not incorporate beliefs. This is a model on which minds are windows onto facts.

When one child systematically passes and another systematically fails a certain kind of false belief task, the difference is not necessarily that one lacks a model of minds and actions. It may be that one child is relying, in this kind of task at least, on a model of minds and actions not incorporating beliefs; whereas the other child relies on a model that does incorporate beliefs.

12.6 The mindreading puzzle

On the face of it, it seems that for many children, there is an age at which:

1 in performing a false belief task of the kind 2- and 3-year-olds tend to fail,709 the child relies on a model of minds and actions not incorporating beliefs; 2 in anticipatory looking and on violation-of-expectation tasks (or in other false belief tasks which 1- or 2-year-olds tend to pass—see Section 13.1), the child relies on a model of minds and actions incorporating beliefs; 3 the child has a single model of minds and actions.

These claims are jointly inconsistent, so one of them must be false. But which claim should we reject? Taken individually, each seems at least superficially plausible. The first claim was discussed in Section 12.4 and the second claim in Section 12.2. And while I haven’t discussed the third claim, rejecting this claim and supposing that children have multiple, incompatible models of minds and actions would surely be a desperate measure.

Solving the Mindreading Puzzle requires us to work out which of these three claims, 1–3, we should reject. In what follows we will consider attempts to solve it.

Notes

tasks (for example, Wellman, Phillips and Rodriguez 2000; Wellman and Liu 2004; Wellman, Fang and Peterson 2011; Rakoczy 2010).

13 Three levels of analysis

This chapter introduces two attempts to solve the Mindreading Puzzle introduced in Chapter 12 (see Section 12.6). The first attempt involves rejecting the claim no. 1 of the Puzzle; and the second two are complementary attempts to reject claim no. 2 of the Puzzle. (The possibility of rejecting claim no. 3 is the topic of Chapter 14.) The hope is that solving the Mindreading Puzzle will take us one step closer to understanding how humans first come to know facts about others’ beliefs, and perhaps about mental states more generally.

13.1 Tracking beliefs without representing them?

Can we reject the second of the three claims behind the Mindreading Puzzle, the claim that younger children rely on a model of minds and action incorporating beliefs? Any attempt to reject this claim will hinge on distinguishing tracking beliefs from having a model of minds and actions which incorporates beliefs. Let me explain.

For a process to track someone’s belief that p is for it to non-accidentally depend in some way on whether she believes that p. Relatedly, to say that someone tracks beliefs is to say that there are processes in her which track some beliefs. By contrast, to say that someone’s responses rely on a model of minds and actions incorporating beliefs implies that making these responses involves identifying certain beliefs as the beliefs of particular individuals and using such identifications in predicting or explaining those individual’s actions. What many experiments actually measure is whether certain subjects can track beliefs: the question is whether changes in what another believes are reflected in the subjects’ anticipatory looking, looking times or predictions about action, say. In principle it is possible to track beliefs without explaining or predicting anything, and without relying on any model of minds and actions.

The distinction between tracking beliefs and relying on a model incorporating beliefs is usually ignored, perhaps because it is so natural to assume that anyone tracking beliefs is likely to be doing so by using a model incorporating beliefs. But we need to recognize that tracking beliefs does not necessarily entail representing them if we are to solve the Mindreading Puzzle by rejecting the second claim—the claim that younger children rely on a model of minds and action incorporating beliefs. And some recent discoveries provide us with an interesting reason not to ignore the possibility that infants are tracking beliefs without relying on any model of minds and actions.

13.2 Altercentric interference

The discoveries concern something called altercentric interference. An adult’s ability to track or report facts about a physical object can be influenced (and sometimes impaired) when another manifestly has a false belief concerning the object. Importantly, this can occur even when the adult is not required to think about the other or her beliefs (van der Wel, Sebanz and Knoblich 2014).

To illustrate, suppose you are observing a scene involving a ball and an occluder. The ball moves around, sometimes stopping behind the occluder and sometimes leaving the scene. At some point the ball is not visible and you are asked to press a button if you think it is behind the screen. This is a simple task which, apparently, tests only your ability to detect an object’s location: no mindreading is required at all. Or so it seems.

Suppose that we complicate the scene slightly by introducing a protagonist who observes some but not all of the ball’s movements. The protagonist never acts at all, but merely observes. Now it can happen that the protagonist is likely to have a false belief about the ball’s location, as she did not observe some of the ball’s movements. Of course this false belief is completely irrelevant to your task, which is just to press a button if you think the ball is behind the screen. And yet it turns out that how quickly you perform your task is influenced by the other’s beliefs.801

This is altercentric interference: your performance is influenced by another’s belief, although their belief is not relevant to your task.

Kovács, Téglás and Endress (2010) provided evidence that altercentric interference occurs in 7-month-old infants and not just adults. Impressively their paradigm allowed them to use the same stimuli with 7-month-old infants and adults.802

What explains altercentric interference? We should distinguish two possibilities. One possibility is that you represent the protagonist’s belief, and this belief representation somehow interferes with your representation of the object. Another possibility is that you do not represent the other’s belief at all. Instead, you believe what they believe.

As the second possibility can be hard to take seriously,803 recall the Motor Theory of Goal Tracking from Section 11.2. Observing another’s actions can cause motor representations in you as if it were you, not her, who was acting (see Chapter 10). And this ‘mirroring’ not only interferes with your own actions (for example, Kilner, Paulignan and Blakemore 2003) but also plays a role in enabling you to predict her actions (for example, Costantini et al. 2014).

Importantly, the Motor Theory does not entail that you have beliefs about, or represent, another’s motor representations. The theory is entirely about first-order motor representations of outcomes.

It is possible that mirroring might occur with respect to representations other than motor representations, including spatial and perceptual representations (for example, Zwickel et al. 2011; Furlanetto et al. 2015; Kampis et al. 2015). And, of course, beliefs.

Observing another person experiencing events that cause her to acquire the belief that the melon is in the box, say, may also cause you to believe this too, and thereby enable you to predict her actions. This is more like contagion than ascription. It not a matter of representing her belief. It requires no understanding of mental states. It is simply a matter of mirroring her beliefs—of believing what she believes.

Mirroring does not involve any model of minds or actions. You can mirror without having the faintest idea of what someone else believes and without any capacity to represent beliefs. Yet mirroring would enable you to track beliefs.

The performance by infants (and adults’) on false belief tasks which measure altercentric interference demonstrates that they can track beliefs. But this performance could in principle be explained by mirroring beliefs, and therefore without postulating any representations of mental states at all. Could the same be true of infants’ performance on other false belief tasks?

Consider Onishi and Baillargeon’s (2005) experiment again. They had infants observe as a protagonist manifestly acquired a false belief. The protagonist then either performed an action which made sense, given the false belief, or else acted as if she had a true belief. Onishi and Baillargeon’s (2005) key finding was that infants looked longer when the protagonist acted as if she had a true belief (see Figure 12.3). This might be because the infants relied on a model of minds and actions to predict the protagonist’s actions. But it might also be because, like Kovács, Téglás and Endress’s (2010) infants, their expectations about the object’s actual location were influenced by the protagonist’s false belief. In this case, the effect would be a consequence of altercentric interference and need not involve infants ascribing a belief or any mental state.

Altercentric interference could in principle explain performance on any experiment in which subjects simply observe a protagonist with a false belief. This includes experiments that rely on anticipatory looking or violation-of-expectation. They do not obviously require mindreading. So experiments with infants using these paradigms do not, by themselves, give us sufficient reason to suppose infants’ responses rely on a model of minds and actions incorporating beliefs. Can we therefore answer the Mindreading Puzzle by rejecting the second of the three claims that give rise to it?804

13.3 Mirroring beliefs?

Things are not quite so simple. False belief tasks designed for older children typically ask questions both about belief and about reality, effectively ruling out the possibility of belief mirroring. You would not get very far asking 1-year-olds questions, of course. But the same can

be achieved by having a task which requires infants to act on the basis of both information about what a protagonist falsely believes and information about what is actually the case. Several tasks measuring how 1-year-olds interact with a false-believing adult impose just this requirement.805

Knudsen and Liszkowski (2012a) claim that 1-year-olds manifest an ability to track beliefs in pointing to warn others of danger. To support this claim, they exploited infants’ tendency to spontaneously point to inform others about things. They put 18-month-olds in a situation where an adult manifestly had a false belief that would probably cause her to reach into a container which, unknown to her, contained something disgusting. They wondered how often infants would anticipate the adults’ action and point to this container to warn the adult with the false belief. To ensure that infants were not motivated to point merely by the presence of an unexpected disgusting thing, Knudsen and Liszkowski set things up so that there were two disgusting things and anyone who lacked the adults’ false belief would be equally likely to accidentally encounter either of them. And in order to isolate the effect of the false belief, Knudsen and Liszkowski compared infants’ pointing responses in another scenario that was as similar as possible except that the adult did not have a false belief specifying the location of an object but instead was completely ignorant. Would infants point more often to one of the disgusting things when the adult falsely believed that an object of hers was at that location? As you can see in Figure 13.1, they did. The bar chart labelled ‘Experiment 1’ shows that infants pointed more often to the disgusting thing that was where the adult falsely believed that her object

Figure 13.1 Eighteen-month-olds manifest abilities to track beliefs in their pointing actions. Both the target and distractor box contain something disgusting. In Experiment 1, an adult falsely believes that something of hers is in the target box; in Experiment 3 the adult is completely ignorant about the location of her object.

Source: Knudsen and Liszkowski (2012a), Figure 2 (part).

was, and the other bar chart (labelled ‘Experiment 3’) shows that they pointed equally to the two disgusting things when the adult was ignorant and had no location-specifying belief. This is evidence that infants’ abilities to track false beliefs enable them to intervene effectively in others’ actions (Knudsen and Liszkowski 2012b provide further evidence).

Belief mirroring could not straightforwardly explain 1-year-olds’ performance on tasks which, like Knudsen and Liszkowski’s, involve interaction. On the face of it, in pointing to warn the adult, infants are combining information about a false belief (which enables them to predict the adult’s action) with information about the actual contents of a container (which is what makes the warning relevant). This indicates that the pointing is not explained by belief mirroring: to appeal to belief mirroring in explaining a process is to suppose that the process fails to distinguish information about the false belief from information about the container’s contents.

What can we conclude? We have been considering whether it is possible to resolve the Mindreading Puzzle by rejecting its second claim on the grounds that infants are mirroring beliefs rather than representing them. For all we know, some belief tracking in infancy (and adulthood too) may be a consequence of belief mirroring. But infants manifest abilities to track beliefs in a wide variety of ways. And the results of experiments involving pointing, helping and other kinds of social interaction make it unlikely that belief mirroring is what explains the full range of 1- and 2-year-olds’ belief tracking. This suggests that at least some of infants’ responses really are underpinned by a model of minds and actions incorporating beliefs, just as the second claim in the Mindreading Puzzle has it.

If you reject this suggestion, you will need to explain why infants manifest belief-tracking abilities on such a wide variety of false belief tasks if it is not because they have a model of minds and actions incorporating beliefs.806 If, on the other hand, the suggestion is right, then we must look elsewhere for a solution to the Mindreading Puzzle.

13.4 Three levels of analysis

If we can’t reject the second of the claims comprising the Mindreading Puzzle, what about rejecting the first claim? According to this first claim, the children we are concerned with fail false belief tasks like Wimmer and Perner’s because, in performing these tasks, they rely on a model of minds and actions not incorporating beliefs. Suppose this claim is false. Why else might these children fail some false belief tasks?

Answering this question involves two steps. We must first identify a feature, or a set of features, common to all the false belief tasks that the children fail. Then, second, we must show that this feature means that these false belief tasks impose an extraneous demand on children. This demand should be extraneous in the sense that it is not related, or only indirectly related, to ascribing false beliefs. Showing this will allow us to conclude that the children fail some false belief tasks not because they rely on a model of minds and actions not incorporating beliefs (as the Mindreading Puzzle has it) but because these tasks involve an extraneous demand.

Pursuing this idea involves switching levels of analysis. Let me explain. Fully understanding mindreading will eventually require at least three levels of analysis (see Table 13.1). First, it requires us to analyse the model or models of minds and actions children and adults use in

Table 13.1 Fully understanding mindreading requires at least three levels of analysis

ModelsWhich model or models of minds and actions do children use in mindreading?
MechanismsHow does a particular model relate to the child’s cognition and action? Is the model something the child knows, does the model describe the operation of a core system, or what?
TasksWhat features determine whether a particular false belief tasks is an A-task?

performing various tasks. This has been our focus until now: the Mindreading Puzzle is primarily a puzzle about which model of minds and actions are used in false belief tasks. Second, it requires us to analyse how the model relates to the child’s or adults’ cognition and action. What kind of mechanism realizes a given model of minds and actions? Is it, for instance, that the model is something known which is explicitly used in inferences? Or should we rather think of the model as describing the operations of a core system, or in some other way? We do not yet have to face up to these questions about mechanisms. Instead our current concern is with a third level of analysis, the task analysis.

What is task analysis? We know that there are at least two kinds of false belief task. There is a kind that children tend to fail until they are somewhere around 3–5 years old, and there is a kind—at least one, maybe more—that children tend to pass in their first or second year of life. We also have good evidence that all the false belief tasks children tend to fail until they are somewhere around 3–5 years old are measuring a single factor (Wellman, Cross and Watson 2001). Let us therefore call any false belief task that children tend to fail until around 3–5 years of age an A-task. I’ll refer to any false belief task that infants tend to pass in their first or second year of life simply as a non-A-task. (I’m tempted to call them ‘B-tasks’ but won’t because we don’t yet know whether different tasks in this category measure different things, as Yott and Poulin-Dubois (2016; Poulin-Dubois and Yott 2017) stress.)

While we have used age to distinguish between A-tasks and non-A-tasks, note that age is not the fundamental issue here. (Researchers sometimes write, mistakenly in my view, as if the developmental puzzle about false belief were that children pass A-tasks relatively late in life.) What really matters is the interaction between performance and task type: there are individuals whose performance on non-A-tasks appears to conflict with their performance on A-tasks. Even if this interaction could be observed for just a day in each child’s life, rather than for several years, its theoretical interest would hardly be less.

13.5 Task analysis

Which features determine whether a particular false belief task is an A-task, a non-A-task, or neither?

An adequate answer to this question should make it possible to determine a priori whether or not a completely novel false belief task is an A-task. That is, the task analysis should enable us to predict in advance of actually testing any subjects how different groups of subjects will perform on the task.

One initially tempting suggestion is that A-tasks are distinct from other false belief tasks in that they involve language in some way. Things are not so simple, however. Some A-tasks are entirely non-verbal (for example, Call and Tomasello 1999), some barely involve language at all (for example, Krachun et al. 2010), and many involve non-verbal responses (for example, Chandler, Fritz and Hala 1989; Custer 1996; Low 2010). Conversely, some non-A-tasks depend on verbal narratives (for example, Scott et al. 2012). A further complication is that a meta-analysis of the relation between false belief and language found no significant difference in the relation between language and belief tracking for different kinds of A-task, despite these differing in the linguistic demands they impose (Milligan, Astington and Dack 2007, 637). This suggests we probably cannot distinguish A-tasks by appeal to the way they involve language.

If it is not simply language, what else could distinguish A-tasks from other false belief tasks? A clue is given by experiments in which two or more kinds of response to the same stimuli were recorded. For instance, recall Clements and Perner’s task (which was introduced in Section 12.6). They showed children a scenario in which someone acquired a false belief and then, for each child, they measured both anticipatory looking and verbal predictions. Low et al. (2014) went even further by measuring anticipatory looking, looking time during a critical window and verbal responses to a question. Two- and 3-year-olds seem to manifest abilities to track false beliefs in their looking times and anticipatory looking but not in responding to a question. This indicates that an A-task and a non-A-task may differ merely in the kind of response measured. But what is it about the response that determines whether a false belief task is an A-task or not?

One hypothesis is that A-tasks involve responses elicited by communicative actions directed to the subject, whereas non-A-tasks do not (Baillargeon, Scott and He 2010; He, Bolz and Baillargeon 2012). A related hypothesis is that A-tasks are those which involve features that somehow disrupt the process of tracking another’s perspective (Rubio-Fernández and Geurts 2012; Rubio-Fernández 2013; Rubio-Fernández and Geurts 2016).

Neither of these hypothesis seems to be right. One objection to both hypotheses is that 18-month-olds appear to perform well on some false belief tasks in which their response is elicited by a communicative action directed to them (for example, Buttelmann, Suhrke and Buttelmann 2015, 98). There is also a converse objection. Performance on A-tasks is related to children’s use of mental state terms in ordinary conversation (Hughes and Dunn 1998), and children appear not to offer unprompted comments on others’ false beliefs until around the age they pass some A-tasks (Bartsch and Wellman 1995; Ruffman, Slade and Crowe 2002; Pyers and Senghas 2009, 810); both facts suggest that A-tasks could in principle include some without disruption and without responses elicited by communicative actions directed to the subject.807

Garnham and Perner (2001, 430) offer a hybrid hypothesis which overcomes these difficulties. Their guess is that a task which involves a ‘declarative expression’ about a belief or belief-based action or emotion will be an A-task;808 a task which involves, or a response prompted by, a communicative action directed to the subject will also be an A-task; and no other tasks will be A-tasks.

Even this hybrid hypothesis is probably untrue. Consider three tasks. In each task, subjects are shown a scenario in which a protagonist manifestly acquires a false belief. In each task,

subjects are required to choose an object. And things are contrived in such a way that subjects have to take into account the protagonist’s false belief if they are to reliably make the best choice. In the first task, subjects are choosing an object for themselves (Krachun et al. 2010). In the second task, subjects are choosing an object to help someone else (Buttelmann, Suhrke and Buttelmann 2015). And in the third task, subjects are choosing an object in accordance with a verbal request (Southgate, Chevallier and Csibra 2010). Which, if any, of these tasks are A-tasks? According to any of the hypotheses we have just been considering, tasks which involve responses prompted by communicative actions directed to the subject should be A-tasks. So all three of these tasks should be A-tasks. In fact, the first is an A-task and the other two are non-A-tasks (see Table 13.2). This suggests that Garnham and Perner’s hypothesis is incorrect. It also illustrates that task analysis is surprisingly hard. It’s not too difficult to make up a story about these three false belief tasks after you know which are A-tasks, of course. The tricky thing is being able to tell the story in advance of knowing which false belief tasks are A-tasks.

My guess is that we are all a long way from understanding what determines whether a false belief task is an A-task. But in the absence of a better supported hypothesis, for now I am going to work on the assumption that Garnham and Perner’s hypothesis is approximately correct. We will eventually consider an explanation for the difficulty of task analysis, one which makes sense of the failure of existing proposals (see Section 14.10).

Why do we care about the task analysis? Our current aim is to solve the Mindreading Puzzle by identifying extraneous demands imposed by A-tasks but not other false belief tasks. In pursuing this aim, we would ideally show that A-tasks are different from other false belief tasks in having some feature (or set of features) which imposes extraneous demands on children, demands that are unrelated, or only indirectly related, to ascribing false beliefs. The point of the task analysis is that it tells us (approximately) which tasks are A-tasks. Our next challenge is to consider whether the distinguishing features of A-tasks do somehow mean that these tasks impose extraneous demands. If they do, it will be possible to solve the Mindreading Puzzle by rejecting the claim that children who can track false beliefs fail A-tasks because they rely on a model of minds and actions not incorporating the possibility of false belief. (This is the first claim in the Mindreading Puzzle.)

Table 13.2 Identifying features that would allow us to predict whether a false belief task is an A-task is hard

TaskType of responseIs an A-task?
Krachun et al. (2010); also Call and Tomasello (1999)Choose an object for oneselfYes
Buttelmann, Suhrke and Buttelmann (2015)Choose an object for someone with a false beliefNo
Southgate, Chevallier and Csibra (2010)Choose an object in accordance with a request from someone with a false beliefNo

Note that some of the findings mentioned here are subject to unsuccessful replication attempts (see Section 12.3) or have yet to be replicated.

It is striking that, in attempting to solve the Mindreading Puzzle, researchers appear to focus on just one level of analysis. Some write as if task analysis alone were sufficient—as if identifying features which distinguish A-tasks from non-A-tasks amounts to answering the questions about models and mechanisms too (compare Rubio-Fernández and Geurts 2012, 6 and He, Bolz and Baillargeon 2012, 26ff). Others appear to neglect the need for a task analysis and so risk mischaracterizing what solving the Mindreading Puzzle requires (for example, Carruthers 2013; Helming, Strickland and Jacob 2015). Both positions are mistaken. Task analysis is both necessary and hard, but a task analysis isn’t a solution to the Mindreading Puzzle all by itself. Solving the Mindreading Puzzle requires not only identifying distinguishing features of A-tasks but also understanding how these features affect subjects’ performance in the ways they do. One false belief task involves a communicative prompt or a declarative expression about belief, whereas another false belief task does not. Why does this apparently minor difference transform the relation between performance on the task and the subject’s age (among other factors)?809

13.6 Selection and inhibition

Take someone and ask her to count some sheep. Suppose she fails. Can you conclude that she can’t count? Maybe, but not if you were whispering random numbers in her ear while she was attempting to count the sheep. In that case, her failure might be a consequence of your distracting ruse rather than of her inability to count. Or suppose that while attempting to count some sheep your subject is also attempting to count the camels and the snakes as well. In this case, too, her failure on your task is not evidence that she cannot count because it could equally be explained by the conflicting demands she is under. She can count alright, she just can’t count three things at once. One premise of the Mindreading Puzzle is that some children fail A-tasks because they cannot count—or, rather, because they are relying on a model of minds and actions not incorporating beliefs. But might it be instead that A-tasks involve extraneous demands? Could the communicative prompt or declarative expression about belief which distinguishes A-tasks somehow function like whispering numbers in someone’s ear while she tries to count? Or could it function like asking her to do something else while counting?

In this section we will investigate the first alternative;810 Section 13.7 explores the second alternative. If either alternative (or some combination of them) is correct, we will have understood why the apparently minor difference between A-tasks and other false belief tasks transforms the relation between performance and age: it is because using a communicative prompt or requiring declarative expression about belief imposes demands extraneous to ascribing beliefs.

Why does this matter? If A-tasks involve an extraneous demand, we have no reason to suppose children who fail A-tasks are relying on a model of minds and actions not involving beliefs. This would mean that we can solve the Mindreading Puzzle by rejecting the first of the three claims that comprise it. (The Mindreading Puzzle was introduced in Section 12.6). This in turn would support a straightforwardly nativist explanation of how humans first come to know facts about beliefs, and perhaps of the mental generally.

The most sustained attempt to show that A-tasks involve demands extraneous to ascribing beliefs is due to Leslie and his collaborators.811 They start from the observation that humans often exhibit a tendency to suppose that others believe what they do. Unless this tendency is inhibited, they are unable to make use of information about false beliefs. And inhibiting this tendency requires mental effort. This means that, even when you as an adult are confronted with a scenario in which it should be obvious that someone has a false belief, you might sometimes act as if they had a true belief (Apperly, Samson and Humphreys 2009, 194).

Extending these ideas, Leslie, German and Polizzi (2005) offer a hypothesis about why many children under 4 or 5 years of age fail false belief tasks. When a child observes a scenario in which someone comes to have a false belief, there is a representation of the content of this false belief in this child; perhaps the content is Maxi’s chocolate is in the blue box. But, in addition, there is also a representation of the content of a corresponding true belief; perhaps this content is Maxi’s chocolate is in the green box. When asked to say where Maxi thinks his chocolate is, the child has to select one of these contents as the content of Maxi’s belief. What determines which content the child will select? According to Leslie, German and Polizzi (2005, 49) this depends on which content has more of a quality they label ‘salience’. Their hypothesis is that the content of the true belief initially has greater salience. According to Leslie et al., this is why younger children and adults under pressure tend to act as if they were unaware of the possibility of false belief. Although they are representing the content of another’s actual false belief, they end up ascribing her a true belief because the content of this true belief is more salient.812

If this is correct, what is involved in ascribing a false belief and passing a false belief task? In order to ascribe a false belief, it is necessary to inhibit the true belief content. This reduces its salience, which means that the false belief content is selected when the child (or adult) ascribes a belief. So why do children under 4 or 5 years of age typically fail false belief tasks? It is because they lack the inhibitory control necessary to ascribe a false belief.

One consequence of Leslie et al.’s view is that whether children can pass false belief tasks should be related to how well they can inhibit tendencies to act. For example, suppose we measured how well children can follow an instruction to avoid peeking at a present or how well they can inhibit a tendency to follow a verbal instruction. Are children who are better at exercising inhibitory control in these cases also better at ascribing false beliefs? It turns out that they are (Carlson, Moses and Claxton 2004; Devine and Hughes 2014), just as Leslie et al.’s view requires.

The evidence is not entirely consistent with Leslie et al.’s view, however. One difficulty is that inhibitory control appears to be only one among several factors which explain whether children pass or fail A-tasks (Devine and Hughes 2014). Further, although factors like inhibitory control can vary between cultures in the sense that two children of the same age from different cultures tend to differ in how well they can exercise inhibitory control, these differences are not reflected in the children’s performance on false belief (Sabbagh et al. 2006). Conversely, deaf children with hearing parents tend to pass A-tasks years later than hearing children (and later than deaf children with deaf parents), and this delay cannot be explained by a corresponding deficit in inhibitory control (Schick et al. 2007; de Villiers and de Villiers 2012).813

A related problem for Leslie et al.’s view arises from studies of the consequences of varying whether a false belief task is likely to require selecting one of several possible contents of belief. In some A-tasks, children are given a scenario in which they know where a protagonist falsely believes an object to be, and in which the children themselves do not know where the object is (for example, Wellman and Woolley 1990). To illustrate, Wimmer and Perner’s false belief task with Maxi and his chocolate introduced in Section 12.1 might be modified so that after Maxi leaves, his chocolate is moved but the child does not know where to. On Leslie et al.’s view, it seems that the need for selection should not arise because the child is in no position to represent the content of a true belief corresponding to the false belief. Despite this, children’s performance on these ‘selection-less’ false belief tasks is not relevantly different from their performance on any other A-task (Wellman and Cross 2001, 704).

A more subtle difficulty for Leslie et al.’s view concerns the role that inhibition plays in ascribing false beliefs. On their view, inhibition is required every time someone ascribes a false belief (without inhibition, the content of the false belief is not selected). But an attempt to directly investigate demands on inhibitory control during belief ascription found no evidence that reasoning about belief specifically required inhibitory control: instead demands on inhibitory control may arise from the need to follow a narrative and answer a question about it, irrespective of whether the narrative concerns mental states (Bull, Phillips and Conway 2008). This suggests that sometimes—perhaps even often—inhibitory control is not required to ascribe beliefs. If that is right, the reason why children do not pass some false belief tasks until they have a certain level of inhibitory control may be that inhibitory control (or something related to it) is necessary for acquiring an ability to think about false beliefs (Benson et al. 2013; Devine and Hughes 2014). This suggests that Leslie et al.’s hypothesis may need refining. I shall assume, for the sake of argument at least, that the view can be refined to accommodate this apparently conflicting evidence.814

There is a more pressing issue. So far I have presented Leslie et al.’s view as if it were an explanation of why children tend to fail false belief tasks until they are around 4 or 5 years of age. This is also how Leslie et al. presented their view before Onishi and Baillargeon’s ground-breaking discoveries about infants’ abilities to track false beliefs (these discoveries were introduced in Section 12.6). But, thanks to these discoveries, we now know that a view which explains why children tend to fail false belief tasks until they are around 4 or 5 years of age would explain too much. What we need to explain is why they fail all and only those false belief tasks which are A-tasks until around this age. How can Leslie et al.’s view be extended to explain this?

Their key idea is that selection is required on A-tasks but not on other false belief tasks (Scott and Baillargeon 2009; Baillargeon, Scott and He 2010; Scott et al. 2010).815 What does this mean? Recall that on A-tasks, subjects are required to ascribe a false belief and this is held to involve selecting between the false belief content (for example, Maxi’s chocolate is in the blue box) and the initially more salient content of a corresponding true belief (for example, Maxi’s chocolate is in the green box). On non-A-tasks, such selection is not required. Or so these authors claim. But why accept this?

In principle, there are two possible reasons why selection would not be required on non-A-tasks. One is that non-A-tasks do not involve ascribing beliefs at all. If this were right, infants’

performance on non-A-tasks would be merely a consequence of the fact that they represent the content of a false belief as well as the content of a corresponding true belief. This might explain the results of false belief tasks in which infants merely distinguish between situations involving true and false beliefs (for example, Kovács, Téglás and Endress 2010; Southgate and Vernetti 2014). But this cannot explain cases in which infants appear to have specific expectations concerning how someone with a false belief will act (for example, Southgate, Senju and Csibra 2007; Knudsen and Liszkowski 2012a). For this reason I take the idea that selection is not involved in non-A-tasks because they do not involve belief ascription to be insufficient by itself.

Another possible reason why selection is not involved (or does not require inhibition) in non-A-tasks is that, in these tasks, infants only represent the content of a false belief and do not also represent the content of a corresponding true belief (Scott and Baillargeon 2009; Baillargeon, Scott and He 2010; Scott et al. 2010). This idea coheres with a more general idea about selection processes: they are needed only because multiple potential targets are represented or attended to (Buetti, Lleras and Moore 2014). The challenge here is to understand why non-A-tasks in particular should trigger a representation of the content of the false belief without also triggering a representation of the content of a corresponding true belief. What is it about non-A-tasks that could explain this? Our attempt at task analysis (in Section 13.5) yielded as a working hypothesis the view that A-tasks differ from non-A-tasks in that A-tasks involve either a declarative expression about belief or belief-based action or emotion or else a response to a communicative prompt. It is unclear how such differences could be responsible for whether those performing the task represent only the content of a false belief or also represent the content of a true belief. This is not to say that it might be possible to develop Leslie et al.’s idea further. But to offer it as an explanation for why some children who can track beliefs nevertheless fail A-tasks amounts to invoking something more deeply mysterious than the thing to be explained ever was.

If you bought into Leslie et al.’s claims about why some children who can track beliefs nevertheless fail A-tasks, I think there’s an excellent chance you can get your money back. Conversely, being able to turn these claims into a genuine explanation should make you famous.

13.7 Too much mindreading?

Our current question is whether we can solve the Mindreading Puzzle by showing that A-tasks impose extraneous demands on children. We have just been exploring the idea that A-tasks involve not just ascribing beliefs but also an extraneous operation of selecting between the possible contents of a belief. As we saw, this idea may eventually be important to understanding the Mindreading Puzzle. But even if it can overcome the various objections, it almost certainly does not identify extraneous demands that are present in all and only those false belief tasks which are A-tasks. Can we do better?

Let us step back from the details to get a sense of the difficulty of our current approach to solving the Mindreading Puzzle. Compare Onishi and Baillargeon’s violation-of-expectation

false belief task with an A-task which, instead of requiring children to predict actions based on false beliefs, instead involves children watching someone perform an action based on a false belief and then requires them to explain this action (for example, Wimmer and Mayringer 1998). Leslie et al.’s view requires that inhibitory control beyond what 2- and 3-year-olds are capable of is needed only for the A-task. But on our current assumptions, performing either task—the non-A-task or the A-task—involves using a model of minds and actions incorporating beliefs to identify a causal relation between a false belief and an observed action. Why might the difficulty involved in identifying such a relation vary dramatically depending on factors such as whether it eventually results in looking longer or in giving a verbal response?

Carruthers (2013, 2015a) has suggested an answer. He proposes that any false belief task which involves a verbal response involves performing three mindreading tasks simultaneously or in close succession. Indeed, he suggests that this is true of all and only A-tasks.816 Mindreading is involved in ascribing belief, in processing a communicative prompt and in producing speech or some other communicative response. So what prevents some children who can ascribe beliefs from manifesting their competence on A-tasks is that they cannot (yet) ascribe beliefs while also doing the mindreading needed to understand and respond to a communicative prompt. Or so Carruthers suggests.

An immediate problem with Carruthers’ suggestion, as I have just presented it, is that it generates incorrect predictions. There are many tasks which are almost identical to an A-task but involve something other than belief; these include tasks about desire, perception, pretence and speech. If Carruthers is right that A-tasks involve performing three mindreading tasks simultaneously, then these related tasks likewise involve performing three mindreading tasks. And if Carruthers is right that children who fail A-tasks despite passing other false belief tasks fail A-tasks because these involve performing three mindreading tasks simultaneously, then we would expect that children should fail other tasks which involve performing three mindreading tasks simultaneously. But they do not. As already mentioned, children who fail A-tasks succeed on almost identical tasks which concern not belief but another mental or semantic state (see Section 12.4).

Since this is not merely a problem for Carruthers but a major obstacle to solving the Mindreading Puzzle, let’s pause to consider some of the tasks. Gopnik and Astington (1988) created a false belief task which is about children’s own beliefs rather than someone else’s. They showed children a Smarties box and asked them what they thought was in it; as expected, most asserted that it contains Smarties. They then opened the box to reveal that it actually contained pencils. Children were then asked what they first thought was in the box. Here is an illustrative protocol from Astington and Gopnik (1988, 195) illustrating how a typical 3-year-old respons:

EXPERIMENTER: Look! Here’s a box. SUBJECT: Smarties! E: Let’s look inside. S: Okay. E: Let’s open it and look inside. S: Oh … holy moly … pencils!

E: Now I’m going to put them back and close it up again. [does so]

E: Now … when you first saw the box, before we opened it, what did you think was inside it?

S: Pencils.

E: Nicky [friend of the subject] hasn’t seen inside this box. When Nicky comes in and sees it … When Nicky sees the box, what will he think is inside it?

S: Pencils.

This is an A-task: children’s performance is correlated with their performance on other variations of Wimmer and Perner’s false belief tasks (Wellman, Cross and Watson 2001). Gopnik and Slaughter (1991) developed a variety of tasks which are structurally identical but concern changes not in what a child believes but in what she pretends, desires, imagines or sees. Strikingly, 3-year-olds who misreport changes in their own beliefs can nevertheless answer similar questions about changes in their own pretence, desires and perceptions. Likewise, Riggs and Robinson (1995) contrasted children’s ability to report what people said with their ability to report what people think. When reporting their own past speech, children who failed to correctly report what they thought also misreported what they said (see also Wimmer and Hartl 1991). But when reporting another’s speech, Riggs and Robinson (1995) found that children who could correctly answer the question ‘What did Iain say?’ failed to answer the question ‘What did Iain think?’

How is all of this relevant to Carruthers’ suggestion about why some children who can track beliefs nevertheless fail A-tasks? If Carruthers is right that A-tasks require performing three mindreading tasks simultaneously, then the counterpart tasks involving not belief but desire, pretence, perception and speech also involve performing three mindreading tasks simultaneously. So children who fail A-tasks can readily perform multiple simultaneous mindreading tasks. The mere fact (if it is a fact) that A-tasks involve multiple simultaneous mindreading tasks therefore cannot explain why children’s performance on these tasks is related to their age in the way it is. Minimally, Carruthers’ suggestion needs supplementing with an account of why multiple mindreading tasks involving beliefs are different from other multiple mindreading tasks.817

Suppose this could be done. Would we have solved the Mindreading Puzzle? Not yet. Carruthers’ suggestion simplifies things by assuming that A-tasks involve both a communicative prompt and a communicative response. In fact, some A-tasks involve just one of these two features. For instance, there are tasks in which you have to respond by selecting an object for yourself (for example, Krachun et al. 2010), or by pointing to a picture (for example, Custer 1996).818 Conversely, consider spontaneous comments on beliefs. When children’s talk about minds is recorded, spontaneously commenting on beliefs is closely related to being able to pass A-tasks (Bartsch and Wellman 1995; Ruffman, Slade and Crowe 2002). So although there are no A-tasks which clearly do not involve a communicative prompt (as far as I know), it seems likely that a sufficiently patient experimenter could devise one. So we must modify Carruthers’ view slightly. His claim must be not that all A-tasks involve three mindreading tasks simultaneously, but that they involve at least two.

This seems like a minor modification but it creates a further problem. Take a look at Table 13.3. Table 13.3 shows that some false belief tasks which are not A-tasks require subjects

Table 13.3 Whether a task involves a communicative prompt or requires a communicative response is not straightforwardly related to whether it is an A-task

Task/paradigmInvolves communicative prompt?Requires communicative response?Is it an A-task?
1 Non- and low-verbal tasks (e.g. Custer 1996; Krachun et al. 2010; Low 2010)YesNoYes
2 Tests by Southgate, Chevallier and Csibra (2010); Buttelmann, Suhrke and Buttelmann (2015); Carpenter, Call and Tomasello (2002)YesNoNo
3 Spontaneous talk about beliefs (e.g. Bartsch and Wellman 1995; Ruffman, Slade and Crowe 2002)NoYesYes-ish*
4 Test by Knudsen and Liszkowski (2012a)NoYesNo

Note that some of the findings mentioned here are subject to unsuccessful replication attempts (see Section 12.3) or have yet to be replicated. (*Children’s spontaneous talk about belief is related to their performance on A-tasks.)

to respond to a communicative prompt (see row 2), and that at least one non-A-task involves subjects producing a communicative response (see row 4). Apparently, then, children who fail A-tasks are nevertheless capable of performing two mindreading tasks where one involves belief.

Here, then, is a second objection to Carruthers’ suggestion. He claims that children fail A-tasks because these tasks involve three mindreading tasks. This is probably false because some A-tasks do not involve three mindreading tasks (at least not for the reasons Carruthers identifies). We might then revise the claim to say that children fail A-tasks because these tasks involve two or three mindreading tasks. But the revised claim conflicts with evidence that children pass some tasks which (according to Carruthers) involve two mindreading tasks.

We have examined two attempts to solve the Mindreading Puzzle by identifying extraneous demands that distinguish A-tasks from false belief tasks. In each case the aim was to provide an alternative to the claim that children fail A-tasks because in performing them they rely on a model of minds and actions not incorporating beliefs. The alternatives are that children fail A-tasks because they require extraneous demands in the form of selecting between multiple possible contents of a belief, or because they require performing multiple mindreading tasks simultaneously. As we have seen, both alternatives face objections. While these objections are not decisive, they do show that the Mindreading Puzzle is more interesting, and harder to solve, than is commonly appreciated. These objections also provide a reason to consider alternative ways of solving the Mindreading Puzzle.

The objections we have considered generalize beyond the specific proposals we have been considering. The broader idea behind Leslie et al.’s and Carruthers’ approaches is that 2- and 3-year-olds who can track false beliefs nevertheless fail A-tasks because of performance factors (compare Leslie, German and Polizzi 2005, 74). As we have seen, a wide variety of findings indicate that 2- and 3-year-olds’ responses on A-tasks are unlikely to be a consequence of performance limits. And correlations between changes in performance on A-tasks

and improvements in things typically considered to be performance factors, such as capacities for selection and inhibitory control, are most likely a consequence of the fact that these capacities are involved in acquiring a model of minds and actions which incorporates the possibility of false belief (as Devine and Hughes 2014 suggest).

13.8 What now?

Our aim is to understand how humans first come to know facts about mental states and beliefs in particular. The obstacle to understanding this is what I have been calling the Mindreading Puzzle (which was introduced in Section 12.6). This puzzle arises because of some apparently conflicting evidence.

There is strong evidence that infants, from sometime in their first or second year of life onwards, can use a model of minds and actions incorporating beliefs in predicting actions and understanding others (as we saw in Section 13.1). Taken in isolation, this evidence would support a straightforwardly nativist answer to our question about the developmental emergence of knowledge of minds. Humans are born with what Kovács, Téglás and Endress (2010) call the ‘social sense’. Their coming to know what someone believes is essentially a matter of directing their social sense at the person.

But there is also strong evidence that typically developing children can first use a model of minds and actions incorporating beliefs in order to predict actions and understand others some time around their fourth or fifth birthday. Acquiring this model of minds and actions appears to take them many months, and to depend on opportunities for discussion about minds and actions as well as on rich forms of social interactions (as we saw in Section 12.1). Taken in isolation, this evidence would support a straightforwardly constructivist account of how humans first come to know about minds. Our understanding of minds is like our understanding of economic value or of astronomy. We acquire it through interactions, particularly interactions with people who, like human adults, are already relative experts, and perhaps also through reflection.

Faced with this apparently conflicting evidence, the obvious strategy for reconciliation is to attempt to discover what is wrong with one or another of these positions. That is what we have been trying to do (and it is what most researchers have been trying to do, too). We have considered whether infants’ performance on some false belief tasks, the non-A-tasks, might be explained without supposing that they are relying on a model of minds and actions incorporating beliefs (see Section 13.1). And we explored whether the patterns in children’s performance on some false belief tasks, A-tasks, might be explained other than by supposing that in performing these tasks they are relying on a model of minds and actions not incorporating beliefs (see Sections 13.5 and 13.7). We have seen that both approaches to reconciling the apparently conflicting evidence and solving the Mindreading Puzzle face major obstacles. My guess is that neither approach will succeed. Even if you don’t agree, the obstacles are substantial enough to justify investigating an alternative strategy.

Notes

Scott and Baillargeon (2009, 1176) write as if response selection and inhibition are distinct processes whereas Leslie, German and Polizzi (2005, 50) clearly regard inhibition as a mechanism by which selection is achieved. My aim here is to reconstruct the most defensible, best-supported version of their view. But their work is so full of ideas that if you read these papers yourself you might end up with a different interpretation of their view.

14 Mind

A solution?

Our aim is to understand how knowledge of minds emerges in development. The obstacle is what I have been calling the Mindreading Puzzle. This puzzle arises from the individual plausibility of three claims which are collectively inconsistent. As we saw in Section 12.6, it seems that for many children, there is an age at which:

1 in performing false belief tasks which are A-tasks, the child relies on a model of minds and actions not incorporating beliefs; 2 in performing false belief tasks which are non-A-tasks, such as tasks involving anticipatory looking or violation-of-expectation, the child relies on a model of minds and actions incorporating beliefs; 3 the child has a single model of minds and actions.

The puzzle is to work out which claim (or claims) to reject. In Chapter 12, I considered the first two claims. For each of these claims, the reasons for accepting them outweigh the reasons for rejecting them. By a process of elimination, we are all but forced to consider rejecting the third and conjecturing that children (and perhaps adults too) have more than one model of minds and actions.

Rejecting the third claim is progress insofar as it allows us to avoid contradiction. But merely rejecting the third claim raises many hard questions. If children have two models of minds and actions, why do they use one in A-tasks and the other in non-A-tasks? What is the nature of these models? Why do they have more than one model? And so on. To gain insight into the developmental origins of knowledge of minds, we must do more than merely reject the third claim about models on the grounds that the first two claims cannot plausibly be rejected. We must identify a theoretically coherent and empirically motivated framework for

thinking about mindreading, one which will allow us to make sense of the possibility that multiple models of the mental coexist in individual humans.

In this chapter we will examine such a framework. We will investigate the idea that there are multiple kinds of mindreading process and compare this with claims about core knowledge. We will also see how to use signature limits in evaluating theories about mindreading. All of this will eventually put us in a position to understand how a combination of core knowledge, social interaction and linguistic structures drives the emergence, in development, of knowledge of minds.

14.1 Mindreading is sometimes automatic

Earlier, when we had to solve a puzzle about infants’ abilities to track briefly unperceived objects (in Chapter 6), we first considered the abilities of adults and then related these back to infants. The same strategy is necessary if we are to understand infants’ abilities to track beliefs.

A key question about adults is whether their mindreading is automatic. The term ‘automatic’ is used in a variety of ways, but throughout this book a process is automatic just if whether or not it occurs is to a significant degree independent of your current task, motivations and intentions. To say that mindreading is automatic is to say that it involves only automatic processes. Is mindreading in adult humans automatic?

One way to show that mindreading is automatic would be to start with a task that does not require tracking beliefs and then to compare subjects’ performance on a measure of belief tracking in two conditions: (1) a condition in which they are not given any prior instructions to track beliefs, and (2) a condition in which they are told in advance that they should track beliefs. Given that the instructions alter subjects’ task, motivations or intentions to track beliefs, if performance on the belief tracking measure did not vary between conditions, we could infer that mindreading occurs somewhat independently of subjects’ tasks, motivations and intentions. That is, we could infer that mindreading is automatic.

Schneider, Nott and Dux (2014) did just this. They showed three groups of adults the same series of videos in which a protagonist happened to acquire a belief about the location of an object. One group was given no instructions, while a second group was told to track the protagonist’s beliefs. For a third group, Schneider, Nott and Dux raised the stakes by instructing participants to keep track of the actual location of a ball in the video, which would be harder to do if they were tracking another’s beliefs. So for the third group, tracking another’s beliefs is not merely irrelevant to their task but may actually hinder performance. Despite this, Schneider, Nott and Dux found evidence in the adults’ looking times that they were all tracking another’s false beliefs irrespective of variations in instructions. This is evidence for automaticity.901

Further evidence that mindreading can occur in adults even when it is counterproductive has been provided by Kovács, Téglás and Endress (2010) and Edwards and Low (2019), who showed that another’s irrelevant beliefs can influence how quickly people press a key, and by van der Wel, Sebanz and Knoblich (2014), who showed that the same belief can influence the

paths people take to reach an object. Taken together, we have substantial (but not conclusive) evidence that mindreading in adult humans sometimes involves automatic processes only (see Schneider, Slaughter and Dux 2017, for a review).

14.2 Mindreading is not always automatic

Does all mindreading in adult humans involve only processes which are automatic? No. It turns out that verbal responses in false belief tasks that are A-tasks are not typically a consequence of automatic belief tracking. To show this, Back and Apperly (2010) instructed people to watch videos in which someone acquires a belief, either true or false, and then, after the video, asked them an unexpected question about the protagonist’s belief (see also Apperly et al. 2006). They measured how long people took to answer this question. Starting with the hypothesis that answering a question about belief involves automatic mindreading only, they reasoned that the mindreading necessary to answer a question about belief will have occurred before the question is even asked. Accordingly, there should be no delay in answering an unexpected question about belief—or, at least, no more delay than in answering unexpected questions about any other facts that are automatically tracked. But they found that people were slower to answer unexpected questions about belief than predicted. Importantly, this was not due to any difficulty with questions about belief as such: when such questions were expected, they were answered just as quickly as other, non-belief questions. It seems that, when asked an unexpected question about another’s belief, people typically need time to work out what the other believes. While not decisive, these findings suggest that not all mindreading is automatic.902

Taking these findings together with the research discussed in Section 14.1, we can conclude that mindreading in adults is probably sometimes automatic and sometimes not. Mindreading that underpins anticipatory looking is often (but not always) automatic. And mindreading that underpins success on an A-task involves processes which are not automatic.

If so, what determines whether a given response will be underpinned by an automatic belief tracking process? This cannot be only a matter of which belief is being tracked—after all, in the experiments we have considered, all the beliefs being tracked are simple beliefs about the location of an object. Instead, whether a response is underpinned by an automatic belief tracking process depends in part on the nature of the response. Anticipatory looking is sometimes based on automatic mindreading (although not always, as we will see below), whereas explicit answers to questions are often based on non-automatic mindreading.

How might these findings about adults help us with the mindreading puzzle, which is about development? Earlier (in Section 12.2), we noted that there is a discrepancy between two measures of belief tracking: if you measure anticipatory looking, it appears that 2- and 3-year-olds can track others’ false beliefs, whereas if you measure verbal responses to a communicative prompt, young children fail to track others’ false beliefs. In this section we have considered adults’ performance on the same two measures—anticipatory looking and verbal responses. As we’ve seen, adults’ anticipatory looking in some false belief tasks is a result

of automatic processes whereas their responses to questions are not. Could these findings be connected? We shall eventually see that they are. But first we need a deeper understanding of mindreading in adults.

14.3 A dual process theory of mindreading

The discovery that mindreading is sometimes but not always automatic motivates considering a dual process theory of mindreading. For our purposes, a dual process theory of mindreading is any theory on which mindreading involves two or more processes which are distinct in this sense: the conditions which influence whether one mindreading process occurs differ from the conditions which influence whether another occurs. Constructing a dual process theory of mindreading is the key to linking findings about adults’ performance with the puzzle of mindreading.

The dual process theory I will consider is simple:

There are two (or more) distinct, independent mindreading processes, one more automatic than the other.

Here two kinds of processes are independent just if the conditions which influence whether a process of one kind will yield an incorrect response differ to some extent from the conditions that influence whether a process of the other kind will yield an incorrect response.

You might say, this is a schematic claim, one lacking any substance. You’d be almost right—and to lack substance is roughly the point of this dual process theory. A key feature of this dual process theory is its theoretical modesty: it involves no a priori commitments concerning the particular characteristics of the processes. Identifying characteristics of the process is a matter of discovery. Further, their characteristics may vary across domains. The characteristics that distinguish processes involved in goal tracking may not entirely overlap with those that distinguish processes involved in segmenting physical objects and representing them as persisting, for example.

There is nothing original about the idea of a dual process theory as such. We have already encountered dual process theories in reflecting on infants’ abilities concerning physical objects (see Chapter 6) and actions (see Chapter 11). Indeed, dual process theories have been offered in several areas (see Evans and Frankish 2009; Sherman, Gawronski and Trope 2014, for some cases). Dual process theories sometimes involve controversial claims about features distinguishing the processes, such as consciousness or speed (see Adolphs et al. 2000). But for our purposes, no such claims are needed. Just as we earlier adopted a lighter account of core knowledge (in Section 8.2), so we can avoid building any assumptions about the nature of distinct mindreading processes into the dual process theory itself. Features that distinguish processes should be discovered, not stipulated.

The dual process theory of mindreading will eventually provide us with a key ingredient we need to solve the mindreading puzzle. But, first, is this theory correct? One way to establish the dual process theory would be to show that automatic and non-automatic mindreading processes are both distinct and independent.

We will identify evidence for independence later, in Section 14.7. What about distinctness? Automatic and non-automatic mindreading processes must be at least somewhat distinct since, by definition, only automatic processes will be to a significant degree independent of task, motivation and intention. To justify accepting a dual process theory, we need evidence of further distinctness. Qureshi, Apperly and Samson (2010) found that automatic and non-automatic mindreading processes are differently influenced by cognitive load, and Todd et al. (submitted) provided evidence that adding time pressure affects non-automatic but not automatic mindreading processes. As we gain more evidence for distinctness, we can be more confident that the simple dual process theory of mindreading is correct.903

In addition to finding evidence for it, there is another, theoretical requirement that must be met before we can accept the dual process theory of mindreading. Why would there be two (or more) distinct and independent mindreading processes in humans? Wouldn’t it be redundant—or, worse, chaotic—to have distinct, independent processes for a single purpose?

14.4 Speed-accuracy trade-offs

Any broadly inferential process must make a trade-off between speed and accuracy. In fact this holds for all animals, not just humans (see Heitz (2014), for a fascinating review). To illustrate, suppose you were required to judge which of two only very slightly different lines was longer. All other things being equal, making a faster judgement would involve being less accurate, and being more accurate would require making a slower judgement. (This idea is due to Henmon (1911), who has been influential although he didn’t actually get to manipulate speed experimentally because of ‘a change of work’ (p. 195).) The same applies to judgements everywhere. Achieving greater accuracy in identifying another’s mental state will require slowing down; identifying mental states more quickly will involve being less accurate. What is the optimal balance between speed and accuracy for mindreading? It varies. In maintaining fluent conversation or rushing to help in an emergency, speed is often paramount. But you might spend the whole night pondering whether Sal really believed Ayesha was away. This is one reason why it would be (or is) valuable to have distinct, independent mindreading processes. Different processes can enable different, and complementary, trade-offs between speed and accuracy.

For an analogy, imagine you are washing dishes with a dishcloth. As the stack of plates grows ever higher, you wash faster and faster. But there are limits to how fast you can reasonably make this process. At some point speeding up further requires a different kind of process. You could exchange the cloth for a long-handled brush, or you could switch from processing dishes sequentially to washing them in parallel. Likewise for mindreading. Having two (or more) distinct, independent processes could enable vastly different trade-offs between speed and accuracy, trade-offs that could not all be achieved with any single kind of process.

Automatic mindreading processes can be fast enough to enable anticipatory looking (for example, Low 2010; Schneider, Nott and Dux 2014) and operate at the speed of action (for example, Edwards and Low 2017), whereas non-automatic mindreading is often relatively slow (for example, van der Wel, Sebanz and Knoblich 2014), and may be error-prone in its

early phases. This motivates considering the possibility that automatic mindreading enables trading accuracy for speed.

But how could accuracy be traded to gain speed in mindreading? In the case of physical cognition, one theory for which there is some evidence holds that different kinds of process achieve complementary speed–accuracy trade-offs by virtue of relying on different models of the physical.904 One kind of process is fast by relying on a relatively simple model of the physical. The simple model enables rapid calculations that are accurate enough in a limited but useful range of situations. Another kind of process trades speed to gain accuracy by relying on a more complex model of the physical (Kozhevnikov and Hegarty 2001; Hubbard 2013, 640). This makes sense. When erecting a washing line or fixing a fence, you can rely on a primitive model of the physical that needn’t require complex calculations. But to land a robot on a comet you’ll need a more sophisticated model of the physical, and this will mean that even simple predictions can require relatively complex calculations. Could different mindreading processes likewise make complementary trade-offs between speed and accuracy by relying on different models of the mental?

14.5 What is a model of minds and actions?

A model of minds and actions is a way mental aspects of the world could be (see Section 12.5). As in the case of the physical, we can use theories of the mental to distinguish models of the mental.

On a widely accepted view, mental states involve subjects having attitudes toward contents (see Figure 14.1). Possible attitudes include believing, wanting, intending and knowing. The content is what distinguishes one belief from all others, or one desire from all others. The content is also what determines whether a belief is true or false, and whether a desire is satisfied or unsatisfied. There are two main tasks in constructing a theory of the mental. The first task is to characterize some attitudes. This typically involves specifying their distinctive functional and normative roles.905 The second task is to find a scheme for specifying the contents of mental states. This typically involves one or another kind of proposition, although some have suggested other abstract entities including map-like representations.906

The canonical theory of the mental features attitudes like belief, desire, knowledge and intention and relies on a system of propositions to distinguish their contents. The relation between a protagonist’s mental states and her actions is primarily a matter of what the contents of those mental states provide reasons for her to do. This theory has been developed

Figure 14.1 Mental states involve subjects having attitudes toward contents

by many philosophers over decades, relying on a mixture of common sense, logic and guesswork.907 The canonical theory of the mental is our best attempt to characterize the model of minds and actions that characterizes adult humans’ most reflective thought and talk about the mind. Although it is largely untested, I shall rely on the conjecture that the canonical model characterizes adult humans’ most reflective thought and talk about minds and actions.908

Could a mindreading process characterized by the canonical model be as fast as automatic mindreading? Could such a process be fast enough to support anticipatory looking, for instance? There are at least two obstacles. First, on the canonical model, mindreading involves using a system of propositions. Propositions are abstract objects in the sense that numbers are. Within limits, you can use sentences to identify propositions in something like the way you can use numerals to identify numbers. Propositions can be arbitrarily connected and nested one within another, and they can be used to mark arbitrarily fine distinctions between possible states of the world. So one obstacle to fast mindreading using the canonical model is that, as standardly understood, mindreading requires the use of a system of abstract objects with complex structures.909

A second obstacle to fast mindreading arises from how mental states are linked to actions on the canonical model. How are predictions about action generated from facts about the agent’s beliefs? On the canonical model, beliefs do not predict actions in isolation. Instead they do so as a whole: which actions can be predicted from a given belief can depend, in arbitrarily complex ways, on anything else that the agent believes (and also on anything the agent desires). In fact, generating predictions about what someone will do from ascriptions of belief involves much the same kind of reasoning that is required to work out what should be done in a counterfactual situation. This is a second obstacle to speedy mindreading: reasoning about what should be done in a counterfactual situation is among the most demanding things humans do.

These are substantial obstacles. If anything should demand time and other scarce cognitive resources, it is surely using a system of complex abstract objects to track states, and deriving predictions by reasoning about what should be done. Given the complexities of the canonical model, it is unsurprising that merely holding simple mental states in mind and using them to make elementary inferences is typically quite time-consuming even for adults (Apperly et al. 2008; Apperly et al. 2011). Understanding how a mindreading process characterized by the canonical model could be as fast as automatic mindreading is a significant challenge.

None of the arguments offered here show that this challenge cannot be met or that the two obstacles are insurmountable, of course. But the absence of any proposal for meeting the challenge motivates considering ways to avoid it. The obvious way to avoid it would be to investigate whether automatic mindreading processes might involve a model of minds and actions other than the canonical model. After all, there is no obvious reason to assume that automatic and non-automatic mindreading processes must both be characterized by a canonical model of minds and actions. Maybe the trade-off between speed and accuracy made by automatic mindreading processes depends on their being characterized by a different model of minds and actions. But what model could this be?

14.6 Minimal models of the mental

In asking how different groups—infants, children, adults, non-human animals—model minds, researchers have almost universally relied exclusively on the canonical theory of the mental. They mostly recognize, of course, that humans typically rely on gradually more sophisticated models of the mental as they grow older. But this is taken to be merely a matter of adding attitudes to the model, so that, say, younger children rely on a model of minds and actions which differs from the adult model in not including attitudes like guessing, supposing and Schadenfreude. What few have yet considered is the possibility that different mindreaders, and perhaps also different kinds of mindreading processes within a single mindreader, might use models of the mental which cannot be specified by anything like the canonical theory. And it is by exploring just this possibility that we will be able to understand how automatic mindreading processes could trade accuracy to gain speed.

The search for simple models of the physical is made easy by the history of science, which contains many theories that are simple, accurate within limits—and wildly wrong. Take impetus theories, for example. These theories say that moving objects have something, impetus, that they gradually lose. When they lose their impetus, they stop moving. If you push them, you impart impetus to them, and that is why they move. Compared to a Newtonian theory, an impetus theory is less accurate. But whereas applying a Newtonian theory involves computing factors like friction and air resistance, an impetus theory rolls these all together into a single thing, namely, impetus. The simplicity of the impetus theory means there is no obstacle to processes characterized by the model of the physical it specifies being fast (Kozhevnikov and Hegarty 2001). We want a counterpart of impetus theories for the mental. We want, that is, a theory which, compared to the canonical theory, sacrifices some accuracy to gain simplicity.

Unfortunately turning to the history of science yields nothing useful on the mental. Fortunately philosophers have provided some simple—but wildly wrong—theories of the mental.

Inspired by these theories, and in particular by my personal favourite, Bennett (1976), we can construct a minimal theory of mind.910 Start by defining an agent’s field as a set of objects. Which objects are in the field changes with time depending on the agent’s location and orientation as well on factors such as lighting conditions, the objects’ movements and acoustic effects, or the presence of barriers. By carefully specifying such factors, we will be able to contrive a definition such that, to a limited but significant extent, the objects in an agent’s field are those the agent can perceive. Equally, by ensuring that the notion of a field is defined in narrowly physical terms, we manifestly avoid complexities associated with perception, such as phenomenology, the possibility of illusion, perceptual modalities, reason-giving and links to other mental states.

Given the notions of a field and of goal-directed action (see Chapter 10 on the latter), we can introduce some mental states, encountering and registration, by specifying their structures and functional roles. Structurally, encountering is a relation between an agent and an object and registration is a relation between an agent, an object and a location. To simplify terminology, let us stipulate that for a registration to be correct is for the specified object to be at the specified location. Functionally:

1 All objects in an agent’s field are encountered by that agent. 2 If an outcome involves a particular object and the agent has not encountered that object, then the outcome cannot be a goal of her actions. 3 If an agent last encountered an object at a location, she registers it as at that location. (And conversely: if an agent registers an object at a location, she last encountered it at that location.) 4 If an outcome involves a particular object, the agent cannot successfully perform an action directed to that outcome unless she correctly registers that object. 5 If an agent performs an action directed to an outcome involving an object, the agent will act as if the object were in the location she registers it in.

In situations where no assignment of encounterings and registrations can make all of the above principles true, later principles outweigh earlier principles.

Although minimal, using this theory of mind would enable you to pass many false belief tasks. To illustrate, recall Wimmer and Perner (1983) false belief task from Section 12.1:

Maxi puts his chocolate in the BLUE box …

This tells us that Maxi registers his chocolate in the blue box (via principles nos 1 and 2).

… and leaves the room to play. While he is away (and cannot see), his mother moves the chocolate from the BLUE box to the GREEN box.

This tells us that Maxi’s registration is incorrect.

Later Maxi returns. He wants his chocolate.

We can now predict that if Maxi attempts to recover his chocolate, he will act as if it were in the blue box (principle no. 5). So representing others’ registrations enables us to track their beliefs.

As it stands, the minimal theory of mind we have just constructed would not enable you to succeed on tasks involving false beliefs about things like colour, shape, function or taste, of course. It is also completely silent on motivational factors, so couldn’t yet be used to succeed on false belief tasks involving desire or preferences. But it is possible to extend the fragment above to overcome these limits while retaining two key features that distinguish a minimal theory of mind from a canonical one. These features are, first, functional roles that can be readily codified; and, second, mental states with simple structures whose contents can be distinguished by things which, like locations, shapes and colours, can be held in mind using some kind of quality space or feature map (using propositions or other complex abstract objects for distinguishing the contents of mental states is not allowed).

Minimal and canonical theories of mind are also markedly different in the range of uses they can be put to. A canonical theory of the mental supports explanation, regulation (of self and others) and story-telling.911 In many ways it is more like a myth-making framework than a

scientific theory. A minimal theory of mind, by contrast, is not designed to be used in any of these ways. It merely enables predictions.

The construction of minimal theory of mind is an attempt to provide a mental counterpart of an impetus theory of the physical. This provides us with a possible explanation of how a mindreading process could trade accuracy to gain speed: it could rely on a minimal, rather than a canonical, model of minds and actions. It will also eventually enable us to remove theoretical obstacles which stand in the way of solving the mindreading puzzle. But before we get to that, we face a more pressing question. How can we discover whether a particular mindreading process uses a canonical or a minimal model of minds and actions? Without an answer to this question, contrasting minimal with canonical models would merely add a theoretical complication.

14.7 Signature limits in mindreading

A signature limit of a model is a set of predictions derivable from the model which are incorrect, and which are not predictions of other models under consideration. We can use signature limits to distinguish competing hypotheses about which model characterizes a process. This approach is well established in the case of physical cognition (as we saw earlier, in Section 6.4), and extends straightforwardly to mindreading. Consider two conflicting hypothesis:

1 Automatic mindreading processes are characterized by a canonical model of minds and actions. 2 Automatic mindreading processes are characterized by a minimal model of minds and actions.

In order to distinguish these hypothesis using the method of signature limits, we first need to identify a prediction generated by minimal models which is incorrect and not a prediction of a canonical theory of mind.

One signature limit on minimal models of the mental concerns false beliefs about numerical identity. These are the kind of false belief Lois Lane has when she falsely believes that Superman and Clark Kent are different people. For the world to be as Lois Lane believes it to be, there would have to be two objects rather than one; her beliefs expand the world. Consider Lois Lane at a time when Clark Kent has disappeared and she is observing Superman performing a daring rescue. Does she know where Clark is? Clearly not. But it is impossible to track her ignorance using a minimal model of minds and actions. This is because where mindreading is characterized by a minimal model, tracking others’ mental states is done using objects themselves rather than senses or concepts or any other kind of proxy. So a minimal mindreader trying to track Lois Lane’s beliefs about Clark Kent does so by representing states, such as encountering and registration,912 which are relations between Lois Lane and Clark Kent. But since Clark Kent is Superman, to encounter Clark Kent is the same thing as to encounter Superman. That is, encountering (or registering) something is like being left

of it. If you are left of Clark Kent then you are also left of Superman (since they are one and the same); likewise for encountering and registration. So when Lois Lane is observing Superman, a minimal mindreader represents Lois as encountering Superman, which is the same thing as representing Clark Kent. A minimal model of minds and actions makes systematically incorrect predictions about anyone who, like Lois, has a false belief about numerical identity. This is one signature limit of minimal models.

The hypothesis that automatic mindreading processes are characterized by a minimal model of minds and actions generates the distinctive prediction that automatic mindreading processes are subject to the signature limits of those models, including the one concerning false beliefs about numerical identity.

Is this prediction correct? Low and colleagues set out to test it (Low and Watts 2013; Wang, Hadi and Low 2015). They created and filmed simplified versions of the Superman story. Their films starred a robot which, like Superman/Clark Kent, looked unexpectedly different on different occasions. The films also featured a protagonist, Lois Lane’s counterpart, who manifestly believed, incorrectly, that this robot was not one thing but two. And much as in the original film the audience (but not Lois Lane) gets to see Clark Kent’s transformation into Superman, so also Low et al.’s films allowed the audience (but not the protagonist) to witness the robot’s transformation. In order to use these films to measure belief-tracking abilities, there was a key moment in the films when it was clear that the protagonist would reach into a box to retrieve a particular robot. At this key moment there were two boxes into which the protagonist might reach. One box actually contained the robot he sought (call this the actual location), whereas the other box was where, given the events of the scenario, the protagonist believed the robot to be (call this the false location). Where would viewers expect the protagonist to reach? If their predictions were based on a canonical model of minds and actions, they should predict that the protagonist will reach for the false location. But if their predictions were based on a minimal model of minds and actions, they should predict that the protagonist will reach for the actual location.

Importantly, Low and colleagues measured subjects’ predictions in two ways, using both anticipatory looking and explicit verbal predictions. Anticipatory looking is likely to be a consequence of automatic mindreading (Schneider, Nott and Dux 2014), whereas explicit verbal prediction are likely to be a consequence of non-automatic mindreading (Back and Apperly 2010). As expected, adult viewers’ verbal predictions were nearly all correct. (That is, they predicted that the protagonist would reach into the false location.) By contrast, adults’ anticipatory looking implied the opposite prediction: at the key moment, most fixated on the actual location. This is a reason to prefer the hypothesis that automatic mindreading processes are characterized by a minimal model of minds and actions.

We should be cautious, of course, in putting too much weight on evidence from a single paradigm. (Indeed, Kulke et al. (2018) have since identified likely confounds which cast doubt on the interpretation of Low and Watts’ (2013) findings.) That said, Low et al. (2014) found the same pattern of results using a different scenario, and Edwards and Low (2017) found evidence for the same signature limit using not only a different scenario but also a different kind of response (reaction times). The variety of evidence significantly strengthens the case. In weighing the evidence, it is also important to note the predictions were formulated before

the experiments to test them were done,913 and that the experimenters were independent of the theorists who formulated the predictions. Even so, all the evidence concerning signature limits on automatic mindreading in adults currently comes from a single lab (Jason Low’s) and it is also all relatively recent. As a general rule, it is prudent to base conclusions on evidence from multiple paradigms and from several labs that has been around for a while. Evidence from other labs, perhaps testing other signature limits of minimal models of the mental, will be informative. Even so, I propose we provisionally accept, on the basis of Low et al.’s various experiments, that automatic mindreading processes are characterized by minimal models of the mental.

This is the key to constructing a theory about mindreading that will enable us to solve the mindreading puzzle and bring us closer to understanding how knowledge of minds emerges in development.

14.8 A developmental theory of mindreading

In constructing a developmental theory of mindreading, the dual process theory of mindreading is a good starting point because it is relatively modest and supported by evidence (as we saw in Sections 14.1–14.3). According to the dual process theory, there are two (or more) distinct, independent mindreading processes, one more automatic than the other.

Until now there was a gap in the evidence for the dual process theory: we had not identified evidence for independence. But we have just seen findings indicating that automatic and non-automatic mindreading processes are not merely distinct (in the above sense) but also independent. That is, the conditions which influence whether an automatic mindreading process will yield an incorrect response are to some extent distinct from the conditions that influence whether a non-automatic mindreading process will yield an incorrect response. How do we know? Recall the findings about the signature limits of automatic but non-automatic mindreading (from Section 14.7). These indicate that, for a single subject responding to a single scenario, automatic mindreading processes can generate an incorrect prediction (or do not generate a prediction) even while non-automatic processes generate a correct prediction. And, conversely, studies with infants in the first three or four years of life indicate the converse: non-automatic mindreading processes can generate an incorrect prediction even while automatic processes generate a correct prediction (see Section 12.2). There seem to be at least two distinct mindreading processes, and these are independent.

But why should there be two (or more) distinct, independent mindreading processes? (This question came up at the end of Section 14.3.) Perhaps it is because they enable complementary trade-offs between speed and accuracy (see Section 14.1). Compared to non-automatic mindreading processes, automatic processes might sacrifice accuracy in order to gain speed. But how could such trade-offs be achieved? In principle, complementary speed–accuracy trade-offs might be achieved by having different kinds of mindreading process rely on different models of the mental (see Section 14.5). Given that non-automatic mindreading processes rely on a canonical model of minds and actions, automatic processes could trade accuracy for speed by relying on a minimal model of minds and actions (see Section 14.6). And the

hypothesis that automatic processes rely on minimal models of the mental combined with facts about the signature limits of minimal models generates testable predictions. So far, attempts to test these predictions on adults have supported the hypothesis (see Section 14.7).

In short, speed–accuracy trade-offs motivate considering a hypothesis about minimal models of the mental, and this in turn generates testable predictions, thanks to signature limits. To the extent that those predictions are confirmed, we can accept the hypothesis about minimal models.914

This hypothesis allows us to extend the dual process theory. There are two (or more) distinct, independent mindreading processes, one or more automatic than the other. And the two processes differ in the kinds of model of minds and actions they rely on: whereas non-automatic mindreading relies on a canonical model of minds and actions, automatic mindreading relies on a minimal model.

Note that even when extended in this way, the theory we are considering remains modest in its ambitions. In terms of Marr’s three levels (see Section 3.1), the canonical and minimal theories of mind provide computational descriptions of automatic and non-automatic mindreading processes. But the theory offered here is silent on algorithms and representations, and on the hardware implementation. A deeper theory with commitments at these levels could generate many additional predictions. But we would need new experimental findings to support (or refute) any such theory. And the theory as it stands is a good starting point for thinking about development.

How can we connect the dual process theory of mindreading with development? Consider a conjecture about development:

In the first three or four years of life, non-automatic mindreading processes do not typically enable belief tracking. What changes over development is typically just that non-automatic mindreading comes to enable belief tracking.

This conjecture is consistent with a limited variety of explanations for infants’ failure on A-tasks. Failure may arise because an A-task measures a response that is not driven by an automatic mindreading process. Alternatively, the response may be a consequence of some combination of automatic mindreading and non-automatic processes, but the non-automatic responses dominate. The conjecture is also consistent with a limited range of different possibilities concerning what changes over development. It may be that non-automatic mindreading processes are initially error-prone in the sense that they yield responses that do not take into account particular mental states, and that they become less error-prone over development. It may also be that the probability that a non-automatic mindreading process occurs increases over development. The conjecture is neutral between these possibilities (or combinations of them).

The conjecture about development links discoveries about mindreading in adults with the mindreading puzzle about children’s development. And it leads to many testable predictions about infants’ mindreading abilities.

One set of predictions arises from the implication that non-A-tasks are tasks on which automatic mindreading processes dominate.915 This is a bold and refutable claim because we

defined non-A-tasks as those false belief tasks that children tend to pass until around 3–5 years of age (see Section 13.5). The claim thus provides a link between infants’ and adults’ performance on false belief tasks. Where a task involves conditions that tend to promote automatic mindreading or to suppress non-automatic mindreading, it should be a non-A-task. And, conversely, any non-A-task should involve conditions that tend to lead to performance in adults being dominated by automatic mindreading processes.

A further set of predictions arises from signature limits on minimal models of the mental. If the conjecture is true, infants’ mindreading involves automatic processes only. Since these involve minimal models of the mental only, infants’ mindreading should be subject to signature limits. In particular, as with automatic mindreading processes in adults, infants’ mindreading should not enable them to track false beliefs about numerical identity.

This prediction has been tested using the same scenarios and measures that, as we saw in Section 14.7, have also been used with human adults. And indeed infants’ performance does appear to resemble automatic mindreading in adults insofar as it is apparently subject to signature limits (Wang et al. 2012; Low and Watts 2013; Low et al. 2014). This prediction about signature limits has also been tested in further experiments using a variety of scenarios and measures in which only children (no adults) participated. The results so far are mixed. Two studies have yielded what appears to be evidence that infants can track false beliefs involving numerical identity (Scott, Richman and Baillargeon 2015; Kampis and Kovács 2016; but see Low et al. 2016, for objections to Scott et al. 2015).916 And two studies have yielded contrary evidence in support of the signature limit (Fizke et al. 2017; Oktay-Gür, Schulz and Rakoczy 2018).

These apparently conflicting findings, together with limited evidence concerning other predictions, indicate that we cannot yet confidently accept or reject the conjecture about development. It may be entirely incorrect, partially correct (perhaps infants’ abilities to track mental states are underpinned by a number of different kinds of processes), or it may even turn out to be correct. Since the balance of evidence is in its favour, I propose we tentatively accept the conjecture about development and rely on it in attempting to solve the mindreading puzzle. Even if this means we will not know that the solution developed here is correct, the solution will still be an improvement on the alternatives, as it is readily testable, already supported by some evidence and not yet known to be incorrect.

Considered alone, there is a theoretical obstacle to accepting the conjecture about development. Why should one mindreading process change during a period of development in which the other is (mostly, at least) unchanging? Combining the conjecture about development with the hypothesis that automatic mindreading relies on a minimal model of minds and actions whereas non-automatic mindreading relies on a canonical model of minds and actions enables us to answer this question. A minimal model of minds and actions—even one rich enough to enable tracking others’ beliefs—can be acquired with relatively little social input. By contrast, the vastly greater sophistication of a canonical model of minds and actions featuring beliefs suggests that acquiring facility with one as a child could require social support for much the reasons that acquiring other sophisticated capacities, such as reading does (compare Heyes and Frith 2014). Given that automatic and non-automatic mindreading processes rely on different models of the mental, minimal and canonical respectively, we can

make sense of the two kinds of mindreading process having quite different developmental trajectories.

We now have a developmental theory of mindreading. There are two (or more) distinct, independent mindreading processes, one more automatic than the other. The relatively automatic mindreading process involves a minimal model of minds and actions, whereas the non-automatic process involves a canonical model of minds and actions. In the first three or four years of life, non-automatic mindreading processes do not typically enable belief tracking. What changes over development is typically just that non-automatic mindreading comes to enable belief tracking.

This theory allows us, finally, to solve the mindreading puzzle.

14.9 How to solve the mindreading puzzle

Recall that the Mindreading Puzzle rests on three claims which are collectively inconsistent. It seems that for many children, there is an age at which:

1 in performing false belief tasks which are A-tasks, the child relies on a model of minds and actions not incorporating beliefs; 2 in performing false belief tasks which are not A-tasks, such as tasks involving anticipatory looking or violation-of-expectation, the child relies on a model of minds and actions incorporating beliefs; 3 the child has a single model of minds and actions.

The puzzle is to work out which claim (or claims) to reject.

Accepting the developmental theory of mindreading just introduced (see Section 14.8) would entail rejecting both claims nos 2 and3.

Claim no. 2 turns out to false because, according to the theory at least, the automatic mindreading in the child relies on a minimal model of minds and actions. This does enables the child to track beliefs in non-A-tasks. But the model features registration, a belief-like state, rather than belief.

Claim no. 3 turns out to be false because both automatic and non-automatic mindreading processes occur in the child, and these involve different models of minds and actions, minimal and canonical. The child passes non-A-tasks because performance on these is dominated by automatic mindreading processes and because automatic mindreading involves a minimal model of minds and actions which enables the child to track others’ beliefs. The child fails A-tasks because performance on these is dominated by non-automatic processes, because non-automatic mindreading involves a canonical model of minds and actions, and because the child’s current canonical model does not feature belief and does not enable her to track others’ beliefs. Instead, it is a model on which what determines how people act is not what they believe to be the case but rather what is the case (see Section 12.4).

As these claims require, the development of non-automatic mindreading appears to involve acquiring facility with a sequence of canonical models of the mental. These models gradually

incorporate a wider range of mental states and a more sophisticated understanding of them. In this way, the canonical model underpinning the child's non-automatic mindreading gradually comes to approximate more closely a full canonical model (compare Wellman, Fang and Peterson 2011; see also Roessler and Perner 2013, for a contrasting but related view). From an adult point of view, over development the child's mindreading becomes gradually less error-prone. It is also possible that in parallel with acquiring a gradually more sophisticated canonical model of minds and actions, children become more likely to engage in non-automatic mindreading. The acquisition of an increasingly sophisticated canonical model of minds and actions over development is plausibly facilitated by developments in linguistic abilities, executive function and social interactions (Low 2010; San Juan and Astington 2012; Devine and Hughes 2014; see Section 12.1).

We have now identified a candidate solution to the mindreading puzzle. But what have we learnt about the emergence in development of knowledge of minds? Before facing this question, let us first tidy up a loose end.

14.10 Task analysis revisited

The mindreading puzzle hinged on an unsatisfactory distinction between two kinds of false belief task. An A-task is one that typically developing children tend to fail until around 3–5 years of age, whereas a non-A-task is one that typically developing children tend to pass in their first or second year of life. Earlier, in Section 13.5, we asked, What determines whether a given false belief task in an A-task, a non-A-task, or neither? An adequate answer to this question should enable us to predict, for a completely new false belief task, which category it falls in. But, as we saw, various existing attempts to distinguish non-A-tasks from A-tasks have all failed. Why is task analysis so difficult?

The developmental theory of mindreading we have been considering provides an answer. On any false belief task, performance is likely to reflect some combination of automatic and non-automatic mindreading processes. And there are multiple ways to turn a non-A-task into an A-task. You can change features of the task so as to increase the probability that non-automatic mindreading will influence performance (call this probability C); for example, you might tell subjects to pay attention to beliefs in advance. Or you can change features of the task so as to decrease the probability that automatic mindreading will influence performance (call this A), perhaps by changing the timing or task demands.917 A further way to turn a non-A-task into an A-task is to change features of the task that influence the probability that automatic mindreading will yield a correct ascription or prediction (call this E_A). You could do some combination of these, of course. Importantly, the effect any given feature has on C, A and E_A is likely to depend on which other features are present.

Existing attempts at task analysis all focus on one or two factors, such as whether a task involves a response elicited by communicative actions directed to the subject. According to the developmental theory of mindreading, the problem of task analysis is the problem of identifying which conditions affect the probability of errors in mindreading, or the probability that automatic or non-automatic mindreading occurs, and of understanding how changes in these

conditions interact. If this is correct, the problem of task analysis has been misunderstood. It is not surprising that prior attempts at task analysis have failed.

There is, of course, a good chance that the developmental theory of mindreading we are considering (see Section 14.8) is incorrect. After all, it makes some readily testable predictions and relatively few of these have so far been tested. It could turn out, for example, that whether a false belief task is an A-task has nothing to do with whether subjects' performance on it is dominated by automatic mindreading processes. In that case, we will need a new theory. But since it has not yet been refuted, we should consider how, other than providing a solution to the mindreading puzzle, the theory might help with understanding the emergence in development of knowledge of minds.

14.11 Is there core knowledge of minds?

We have seen that there are (at least) two distinct, independent kinds of mindreading processes in humans, one more automatic than the other. The relatively automatic process enables belief-tracking from early infancy, certainly from early in the second year of life, perhaps from 6 months of age (Southgate and Vernetti 2014), and possibly even earlier. Automatic mindreading in infants also exhibits some features associated with the operations of core systems. As far as we know, it is largely unchanging over the course of development (see Low and Watts 2013; Low et al. 2014; Edwards and Low 2017, for evidence that a signature limit persists over development). Given that automatic and non-automatic mindreading are independent processes, automatic mindreading is also likely to be informationally encapsulated to some degree. These considerations justify provisionally concluding that infants' earliest mindreading abilities do not rest on knowledge of others' mental states. Given the lighter account of core knowledge (introduced in Section 8.2), we may conclude that infants have core knowledge of minds.918

Concluding that infants have core knowledge of minds raises more questions than it answers. In the case of physical objects, we were able to identify constituents of core knowledge: object indexes, motor representations and metacognitive feelings (see Chapters 6 and 7). For each of these constituents, there are reasons to believe it exists independently of any theory invoking core knowledge. A future challenge for proponents of core knowledge of minds is to identify its constituents.

14.12 Origins of knowledge of mind: rediscovery

How does knowledge of minds emerge in development? On the view considered and partially defended in this chapter, there are two (at least) distinct, independent mindreading processes, one emerging early in infancy and another first enabling humans to track false beliefs from around 4 years of age. The former, early-developing kind of mindreading is relatively automatic and appears to involve core knowledge of mind. By contrast, the second, non-automatic kind of mindreading is necessary for knowledge of others' minds. So, as in the

case of physical objects and colour, to understand how knowledge of minds emerges in development, we need to ask what the role of core knowledge is. What is the role of early-developing, relatively automatic mindreading in explaining the emergence in development of knowledge of minds?

Some studies have begun to address this question by looking for correlations between infants' success on early mindreading tasks and the subsequent emergence of abilities to succeed on A-tasks. One study found no correlation between these (Grosse Wiesmann et al. 2016), while another study did find a correlation using slightly different measures (Low 2010). Importantly, both studies found that performance on A-tasks, which measure knowledge of mind, was correlated with factors that do not appear to influence infants' success on early mindreading tasks. These include components of linguistic mastery and inhibitory control.

The emergence of knowledge of minds in development appears to be a consequence not only of early-developing mindreading capacities but also of developments in linguistic abilities, executive function and social interactions (see Section 12.1 and especially San Juan and Astington 2012; Devine and Hughes 2014). The importance of language and social interaction in acquiring a full canonical model of minds and actions is illustrated by cases in which these are deficient. Some adults who lack a syntactically sophisticated language may never acquire a full canonical model of minds and actions.919 And individuals born deaf into hearing families do not pass A-tasks until years later than hearing individuals (Peterson and Siegal 2000; Peterson 2009), suggesting that their more limited opportunities for rich social interactions may make acquiring a full canonical model of minds and actions harder.920

Acquiring knowledge of minds appears to be another case in which development is rediscovery. There is an early-developing capacity to represent mental states, but this does not appear to be connected in any straightforward way to later-developing knowledge of mental states. Instead the emergence of this knowledge hinges on social (and cognitive) skills.

This picture is complicated by the fact that an early-developing capacity to represent mental states plausibly enhances a child's social skills, and may also influence her exposure to situations that promote the development of linguistic skills. Perhaps an early-developing form of mindreading makes possible social interactions which, together with linguistic and cognitive developments, eventually enable us to know facts about others' minds.

Notes

15 Joint action

The overarching challenge we face is to understand the developmental emergence of knowledge. How do humans first come to know simple facts about objects, actions and minds? So far we have seen that infants in the first year or two of life have significant abilities in each domain, and, more controversially, that these abilities do not appear to require attributing them any knowledge states at all. In exploring evidence on development in several domains, we found no compelling reason to accept that 1-year-olds have knowledge, rather than merely core knowledge, in these domains. Instead their abilities rest on what we are loosely calling core knowledge (see Section 8.2), which, at least in the domains of objects and actions and perhaps even in the domain of minds too, comprises broadly perceptual and motoric representations.

These discoveries provide us with an opportunity and a challenge. They make it theoretically coherent to suppose that infants' earliest abilities provide a basis for the later acquisition of knowledge. So knowledge does not have to come from nothing: there is a representational precursor which manifests itself early in development and presumably provides some kind of basis for the later acquisition of knowledge. But—here is the challenge—the discoveries also make it difficult to understand how there could be any connection between core knowledge and knowledge proper. This is because in each domain—objects, actions and minds—we have seen that the representations comprising core knowledge are not inferentially integrated with knowledge states, and that they are intentionally isolated from them. There is little prospect, then, that the transition from not knowing any simple facts about particular objects, actions or minds to knowing some such facts will involve somehow transforming core knowledge states into knowledge states. The developmental emergence of knowledge must be a process of rediscovery: of coming to know things which are in some sense already encoded in core knowledge.

How could rediscovery occur? Perhaps it rests on social interactions. One possibility is that the emergence of knowledge is linked to the appearance of increasingly rich forms of social

interaction in the second year of life. Could such rich forms of social interaction somehow facilitate the developmental emergence of knowledge?

According to what Moll and Tomasello (2007) call the 'Vygotskian Intelligence Hypothesis', 'participation in cooperative … interactions … leads children to construct uniquely powerful forms of cognitive representation'. Knoblich and Sebanz arrive a similar claim based on very different considerations: 'perception, action, and cognition are grounded in social interaction' (2006, 103). And Sinigaglia and Sparaci (2008) likewise argue that: 'human cognitive abilities … [are] built upon social interaction'.

I assume that they are all right, although I am less sure that I understand precisely what each duo of theorists is claiming. To avoid a tricky exegetical detour, let us make some simplifying assumptions. What forms of 'cognitive representation' should these theorists have in mind? I shall assume that it is knowledge. And what form of 'cooperative' or 'social interaction' should they have in mind? I think it should be joint action. These assumptions yield what I shall call the Joint Action Conjecture:

Abilities to perform joint actions play a role in explaining the developmental emergence of knowledge, including knowledge of others' minds.

This is probably not exactly what any of the three theoretical duos quoted above had in mind, but it is inspired by their ideas and findings.

In this chapter, our primary aim is to characterize joint actions as performed by children in the first years of life. If the Joint Action Conjecture is right, this will be a step towards understanding the developmental emergence of knowledge. But first, what are joint actions?

15.1 Joint action vs parallel but merely individual actions

I have been using the term 'joint action' without explanation. What anchors our thinking about joint action? Researchers often introduce it just by giving examples. Paradigm cases in philosophy include two people painting a house together (Bratman 1992), lifting a heavy sofa together (Velleman 1997), preparing a hollandaise sauce together (Searle 1990), going to Chicago together (Kutz 2000), and walking together (Gilbert 1990). In developmental psychology, paradigm cases of joint action include two people tidying up the toys together (Behne, Carpenter and Tomasello 2005), cooperatively pulling handles in sequence to make a dog-puppet sing (Brownell, Ramani and Zerwas 2006), and bouncing a block on a large trampoline together (Tomasello and Carpenter 2007). Other paradigm cases from research in cognitive psychology include two people lifting a two-handled basket (Knoblich and Sebanz 2008), putting a stick through a ring (Ramenzoni et al. 2011), and swinging their legs in phase (Schmidt and Richardson 2008, 284). We need to be careful because it is not obvious that there is a single phenomenon of which all these are paradigm cases.

A slightly better way to introduce joint action is by using contrast cases. Contrast cases are pairs of events which are as similar as possible except that one is a joint action while the other is not. Gilbert (1990) contrasts two friends out walking together in the way friends

typically walk together with two strangers who happen to be walking side by side. The friends' and the strangers' movements may happen be so similar that if looking at them from above, you could not say which pair was the friends and which the strangers. Despite this, the friends' walking together is a joint action whereas merely walking side by side is not. Relatedly, Searle (1990) contrasts an event involving several park visitors simultaneously running to a central shelter in order to perform a dance with an event involving park visitors running to the central shelter in order to escape a storm. The first event is a joint action whereas the second is not. But we could imagine that the two events involve very similar movements.

These contrast cases invite the question, How do joint actions differ from actions that are performed in parallel but are merely individual actions? Minimally, an account of joint action needs to provide an answer to this question.

Gilbert's contrast case shows that the difference is not just a matter of coordination. After all, strangers who find themselves walking side by side may need to coordinate their actions in order to avoid colliding. And Searle's contrast case shows that the difference between joint action and parallel but merely individual actions is not just that the actions have a common effect because parallel but merely individual actions can have common effects too. So what does distinguish joint actions from parallel but merely individual actions?

15.2 Shared intention

Many philosophers and some psychologists hold that all joint actions involve shared intention. They tend to assume that characterizing joint action is fundamentally a matter of characterizing shared intention. On this view, 'the key property of joint action lies … in the participants' having a … "shared" intention' (Alonso 2009, 444–5). Although their terminology differs, many philosophers including Gilbert (2006, 5) endorse this claim, as do developmental psychologists who have thought deeply about joint action, such as Carpenter (2009, 381), Rakoczy (2006, 117) and Tomasello (2008, 181).

But what is shared intention?

The first thing to note is that 'shared intention' is a term of art. On almost any account, a shared intention is neither something literally shared nor literally an intention.

To see the point of invoking shared intention, whatever exactly it turns out to be, it is helpful to draw a parallel with individual action. Davidson (1971) opened a famous discussion of agency by asking, Which events are actions? He contrasted actions with things that merely happen to an agent. To illustrate, we might be struck by the contrast between mere reflexes, such as the eyeblink reflex, and actions such as the blinking your eyes performed in covertly greeting a friend.

One quite standard way to answer the question, Which events are actions? is by appeal to intention. The idea is that events are actions in virtue of being appropriately related to an intention. On this view, intention is what distinguishes a reflex from an action. Intention is absent from, or not appropriately related to, the eyeblink reflex but (so the view) is critical for the action of blinking your eyes.

The question, Which events are actions? has its counterpart for joint actions, namely, Which events are joint actions? Invoking shared intention makes it possible to give an answer to the question about joint action that resembles the answer standardly given about action: a joint action is an event which is appropriately related to a shared intention (Pacherie 2013, 3–7). This parallel between intention and shared intention is important because it is a rare point on which almost everyone will agree—even those who, like me, are not persuaded by the standard view of actions and intentions. Shared intention is to joint action at least approximately what ordinary, individual intention is to ordinary, individual action.

It is important to acknowledge that we have not yet said anything very informative about what shared intention is. We asked, Which events are joint actions? The answer was, those which stand in an appropriate relation to a shared intention. Now we ask, What is shared intention? Suppose we answer by saying that it is something in virtue of which events are joint actions. Maybe this is right. But at this point we have moved in a circle too tight to count as sufficiently informative, even to a philosopher.

Beyond mostly accepting that shared intention is to joint action what intention is to action, philosophers disagree quite wildly about what kind of thing shared intention is. Some, like Searle (1990), regard it as a special kind of mental state, something that differs from intention as belief differs from desire. Others, like Helm (2008; Laurence 2011), appear to think that it requires us to postulate a special kind of agent. Still others, like Gold and Sugden (2007), think that a shared intention is just an ordinary intention that is arrived at by a special kind of reasoning: team reasoning. Taking yet another line, Gilbert (2013) proposes that we understand shared intention in terms of a special kind of commitment, and Tuomela (2005) that we understand it in terms of a special mode: the 'we-mode'. Meanwhile, Bratman (2014), following a strategy introduced by Tuomela and Miller (1988), argues that shared intention does not require any such conceptual or ontological novelty but can be realized by a nested structure of intention and knowledge. It seems as if each philosopher who has thought about this has ended up with a new account of her own, one that is fundamentally different from everyone else's.

So which account of joint action should we start with? Bratman's is probably the most carefully developed and certainly the most influential among scientists.

15.3 Bratman on shared intention

We aim to understand joint actions as performed by 1- and 2-year-olds. And we are assuming, for now at least, that shared intention is the distinguishing feature of all joint actions. We therefore need an account of shared intention.

In characterizing shared intention, Bratman first identifies its function. On his account, shared intention serves to coordinate activities, coordinate planning and structure bargaining (Bratman 1993). To illustrate, suppose you and I have a shared intention that we wash the glasses together. Then this shared intention should structure our bargaining insofar as our different preferences may mean we need to compromise over how clean we will make the glasses; it should coordinate our planning insofar as we need to order our washing and drying efficiently; and it should also coordinate our activities when, for instance, you are passing me a glass to dry.

If this is what shared intentions are for, what could they be? Bratman argues that the following are collectively sufficient conditions for you and I to have a shared intention that we J:

  1. (a) I intend that we J and (b) you intend that we J
  2. I intend that we J in accordance with and because of 1a, 1b, and meshing subplans of 1a and 1b; you intend that we J in accordance with and because of 1a, 1b, and meshing subplans of 1a and 1b
  3. 1 and 2 are common knowledge between us.

(Bratman 1993, View 4)

To illustrate, consider our shared intention that we wash the glasses together again. According to the above account, to have such a shared intention, it would be sufficient that we each intend that we wash the glasses together, that we each intend to wash the glasses together in accordance with and because of these intentions (and meshing subplans of them), and that all of this is common knowledge between us.

What does Bratman mean by subplans? Our shared intention to wash the glasses together may require us to make further plans. For example, we may plan to do the water glasses before the beer glasses, and we may plan to soak the wine glasses before washing them. These plans are subplans just because they are plans that are components of our larger plan to wash the glasses. And the subplans mesh insofar as we can successfully execute all of them—they would fail to mesh if, for example, your subplan was to soak the wine glasses while washing the water glasses while my subplan was to wash the wine glasses first.

In more recent work Bratman has added these further conditions:

  1. The persistence of each intention in conditions 1 and 2 is interdependent with the persistence of every other such intention (Bratman 1997, 153; Bratman 2006, 7–8; Bratman 2009, 157; Bratman 2010, 12)
  2. We will J 'if but only if 1a and 1b'.

(Bratman 1997, 153; 2009, 157)

The common knowledge condition, no. 3 above, is extended to include these further conditions, nos 4 and 5.

You could spend a long and probably very happy time studying the intricacies of Bratman's account and various critical exchanges over it (for example, Bratman 2015; Ludwig 2015; Smith 2015). But this turns out not to be necessary for our purposes. There is a straightforward and compelling objection to combining Bratman's view with the Joint Action Conjecture.

15.4 An inconsistent triad

To meet the sufficient conditions Bratman gives for having a shared intention (see Section 15.3), it is not enough that we each have intentions. In addition, we must each have intentions about these intentions. And, thanks to the common knowledge condition, we must each know that we know that we have intentions about these intentions (see Figure 15.1).

Figure 15.1 On Bratman's sufficient conditions for shared intention, sharing is escalating. Minimally, we each have to know that we know that we have intentions about our intentions.

Figure 15.1

Suppose for a moment that all joint action involved shared intention, and that meeting Bratman's sufficient conditions was the only way to have a shared intention. Then performing joint actions would require having knowledge of others' minds at close to the limits of adults' capacities. In that case we would have to reject the Joint Action Conjecture, on which abilities to perform joint actions play a role in explaining the developmental emergence of knowledge, including knowledge of others' minds. After all, if performing joint actions already requires having knowledge of others' minds, we can hardly invoke abilities to perform joint actions to explain the emergence of such knowledge.

Nothing that involves meeting the sufficient conditions given by Bratman could explain the developmental emergence of knowledge of others' minds because meeting these conditions requires already having knowledge of others' minds.

One way around this issue would be to find alternative sufficient conditions for shared intention, conditions you can meet without already having knowledge of others' minds—or any of the other knowledge we might hope to explain by invoking joint action. Several philosophers have offered ideas that take us at least part of the way in this direction (for example, Tollefsen 2005; Pacherie 2013). But I think there is a reason why this strategy is unlikely to work.

One of the functional roles of shared intention is to coordinate planning, which is distinct from coordinating activities. Planning occurs, for example, when we consider whether to wash the glasses in any particular order. Although there are probably cases in which you can coordinate your plans with another without knowledge of her intentions and other mental states, capacities to coordinate planning will in general require knowledge of others' minds (Butterfill 2012b). Abilities to coordinate plans with another presuppose abilities to know things about others' minds. The following three claims are therefore collectively inconsistent:

  1. Abilities to perform joint actions play a role in explaining the developmental emergence of knowledge, including knowledge of others' minds. (This is the Joint Action Conjecture.)
  2. All joint action involves shared intention.
  3. A function of shared intention is to coordinate two or more agents' plans (as Bratman's account implies).

These three claims cannot all be true. But which should we reject?

To answer this question, we need to consider the capacities of 1- and 2-year-olds. Are children of this age already capable of coordinating their plans with the plans of others? If so, we may have grounds for rejecting the Joint Action Conjecture; if not, there may be reason to reject one of the other claims and find an alternative way to characterize joint actions as performed by 1- and 2-year-olds.

Although the inconsistent triad above is not explicitly about knowledge, the underlying issue is one of knowledge. In fact, the inconsistent triad is analogous to the puzzles we have encountered in the attempt to understand developments about mindreading, action tracking and object cognition in earlier chapters. In each case, there seem to be reasons to suppose that explaining infants' early-developing abilities requires attributing them knowledge of a domain. If decisive, these reasons would make it theoretically incoherent to appeal to these early-developing abilities in explaining the developmental emergence of knowledge in that domain. But so far those reasons have not turned out to be decisive; there are as strong, or stronger, reasons not to attribute knowledge in explaining the early-developing capacities (see Chapter 4, for example). In the case of joint action, coordinating plans requires knowledge of other's mental states, as we have just seen. This is why the question of whether 1- and 2-year-olds' joint actions really involve coordinating plans matters in explaining the developmental origins of knowledge. If these children's actions involve coordinating plans, then they involve knowledge of mental states, and so we would have to abandon the Joint Action Conjecture and avoid appeal to joint action in explaining the developmental emergence of knowledge.

15.5 Coordinating planning

Carpenter argues that infants' joint actions are not less sophisticated in kind than those of adults:

Despite the common impression that joint action needs to be dumbed down for infants due to their 'lack of a robust theory of mind' … all the important social-cognitive building blocks for joint action appear to be in place: 1-year-old infants understand quite a bit about others' goals and intentions and what knowledge they share with others.

(Carpenter 2009, 383)

On her view, which appears to be quite widely shared (for example, Tomasello et al. 2005; Moll and Tomasello 2007), infants are capable of joint action and all joint action involves shared intention as characterized by Bratman. If so, we must reject the Joint Action Conjecture. But first let us consider how well the evidence supports Carpenter's position. The hypothesis that 1- and 2-year-olds have shared intentions as characterized by Bratman generates a prediction: since a function of shared intention is to coordinate planning, children of this age should be capable, at least in some minimally demanding situations, of coordinating their plans with those of another.

What might count as evidence that you can coordinate your plans with those of another? Paulus (2016) created an elegantly simple task which demands minimal coordination. Their apparatus involved a box with two holes, one round and the other square, and a tool with two ends, one spherical and the other cubic. Inserting the tool into the round hole has one effect; inserting it into the square hole another. Which hole the tool should be inserted into depends on your task. And of course the tool will only go into the holes if it is inserted the right way around—the spherical end will not go into the square hole. Now imagine picking up the tool and putting it into the square hole. How should you best pick up the tool in doing this? If you pick it up by the spherical end, you can insert the tool directly into the hole. But if you pick the tool up by grasping the cubic end, you will have to swap it between your hands before you can insert it. This is almost too straightforward to mention: even 3-year-olds pick up the tool optimally depending on the task (Paulus 2016). But now imagine doing the task with another person. You will pick up the tool and pass it to her, and she will insert it into the square hole. How should you best pick up the tool this time? By grasping the cubic end and passing the tool to her, she will take the spherical end and so be poised to insert the cubic end directly into the hole. And this is what adults tend to do. But what about children?

Paulus (2016) found that, when acting jointly with another, 3- and 5-year-olds picked up the tool without any regard for what the other person was going to do with it. It was only 7-year-olds who took into account the other's action, and they only did so after the first few trials. Apparently it is not until surprisingly late that children can coordinate their plans with another's even in this minimally demanding situation.

We should not read too much into a single experiment, of course. It is possible that some extraneous feature unrelated to demands on coordination of plans threw the younger children off in Paulus's (2016) experiment. But it turns out that the developmental pattern is robust across many experiments. Apparently children under about 5 years of age do not coordinate their plans with those of others.

Warneken et al. (2014) created a situation in which a pair of children needed two different tools to release a ball: a plunger and a twister. Before the children could attempt to release the ball, first, one child selected a tool and then the other child selected a tool. How children

selected tools was therefore a good indicator of their planning. In one condition, the task of choosing a tool was made relatively easy. One child chose either a plunger or a twister first, and then the second child could choose whether to take a plunger or a twister. All she needed to do was to take the kind of tool not yet chosen; otherwise they would end up with two tools of the same kind and so be unable to release the ball. The second child's successfully choosing a tool complementary to the one already chosen would indicate an ability to coordinate actions. In another condition, the task of choosing a tool more clearly required coordinated planning. The first child could see that the only tool available to the second child would be a twister. She therefore had to anticipate this constraint on the second child's future choice by taking a twister herself. Success in doing this might be taken to indicate an ability to coordinate planning. Whereas the 5-year-olds did well in both conditions, the 3-year-olds showed no evidence at all of coordinating their plans with the those of another. Some did learn over the course of the experiment to select a tool complementary to the one already chosen, suggesting an ability to coordinate actions. But none showed evidence of coordinating plans by anticipating the constraint on the second child's future choice.

The systematic failure of 3-year-olds to coordinate their plans with those of others suggests that there is a mismatch between our current assumptions about what joint action is and these (and younger) children's actual abilities to engage in joint action. On our current assumptions, all joint action involves shared intention and Bratman is right about what shared intention is. These assumptions are worth taking seriously because they underpin views held by Carpenter (2009) and Moll and Tomasello (2007) among others. But they also appear to be wrong. For Bratman's is a planning theory of shared intentions: on his theory, shared intentions serve to coordinate plans. The hypothesis that 1- and 2-year-olds have shared intentions therefore leads to the prediction that they should be capable, in some situations at least, of coordinating their plans with those of others. And all the available evidence suggests this prediction is false. In fact, coordinated planning is something humans first achieve some time after their fifth birthday.

You might object that the experiments mentioned so far all involve tools. Could there be something especially difficult about coordinating planning when tools are involved? If there is, it is unlikely to explain apparent deficits in coordinating planning because children show the same deficits even when no tools are involved. Meyer, Wel and Hunnius (2016) created a simple scenario in which children passed objects to an adult. The twist was that the children could not pass the objects straight to the adult because of a glass screen between them. Instead children had to pass objects to the adult's left or right. The adult always had only one hand free to receive the object, and which hand she had free changed as the task unfolded. To avoid the adult having to make an awkward reach, children would ideally always pass objects to the adult's free hand, switching sides as necessary. They were clearly to some extent sensitive to this constraint because they mostly began by passing the object to the adult's free hand. But over time, children who were 3 years of age or younger did not pass objects to the adult's free hand, and even 5-year-olds 'adjusted their action plans to a surprisingly small degree' (Meyer, Wel and Hunnius 2016, 8). This suggests that although children under 3 years of age (and perhaps older children too) can anticipate and respond to others' actions, they cannot coordinate planning with them.

To be sure that the evidence we have been considering really does point to a deficit in coordinating planning, it would be helpful to compare two scenarios which are as similar as possible except that one does, whereas the other does not, require coordinating plans.

Gerson, Bekkering and Hunnius (2016, experiment 2) created just such a pair of contrasting scenarios. Three-year-olds were given four egg cups, two light and two dark. They then had to retrieve each of four balls from a transparent tube and place it in an egg cup of the corresponding colour (see Figure 15.2). One ball was half-dark and half-light and was allowed to be in any egg cup. But since children only had four egg cups, they needed to pay attention to the colours of the remaining balls in placing the half-dark, half-light ball. Otherwise they would not have enough egg cups of the correct colours for the remaining balls. When performing the task alone, children did well enough. But Gerson, Bekkering and Hunnius (2016) were most interested in what would happen when children performed the task jointly with another (a puppet, in this case), taking turns to place balls. They found that children were much worse when acting jointly. In fact they showed no evidence at all of planning ahead when acting jointly (Gerson, Bekkering and Hunnius 2016, experiment 1). But most importantly, the researchers compared the joint action scenario with a machine scenario. In the machine scenario, children alternated with a machine which placed balls into egg cups. The machine scenario was as similar as possible to the joint action scenario except that no coordinated planning was needed (or even possible). Despite this, 3-year-old children's performance was better in the machine scenario and no different from their performance when they were performing the whole task alone (Gerson, Bekkering and Hunnius 2016, experiment 2). Together

Figure 15.2 Take each ball from the bottom of the tube and put it in a matching egg cup. Schematic representation of the task in Gerson, Bekkering and Hunnius (2016).

Figure 15.2

with the other findings we have considered, this is strong evidence that 3-year-olds cannot yet coordinate their plans with those of others.

We are working on the assumption that a function of shared intention is to coordinate two or more agents' plans. Given this assumption, the hypothesis that 1- and 2-year-old children have shared intentions leads to the prediction that these children can coordinate their plans with those of others. At least, they should be able to do so in minimally demanding situations. But in fact it appears that abilities to coordinate plans develop much later, perhaps between 5 and 7 years of age. It is always possible, of course, that the available evidence is misleading and a new breakthrough finding will significantly alter the picture. But the available evidence clearly best supports the hypothesis that 1- and 2-year-old children are not capable of having shared intentions.

15.6 Joint action in the first years of life

We can summarize the position we have reached so far in this chapter with another inconsistent triad (the first was in Section 15.4):

  1. One- and 2-year-olds are capable of performing joint actions.
  2. All joint action involves shared intention.
  3. A function of shared intention is to coordinate two or more agents' plans (as Bratman's account implies).

As we saw, Carpenter and others hold that all three claims are true. But these claims lead to the incorrect prediction that 1- and 2-year-olds are capable of coordinating their plans with those of others. For this reason, at least one of the claims should be rejected. But which?

Consider the first claim, which is that 1- and 2-year-olds are capable of performing joint actions. So far, we have been taking the truth of this claim for granted. But a hard-line response to the evidence that children show no sign of coordinating their plans with those of others until around 5 years of age would be to deny that children much under this age are capable of joint action at all. Are there grounds to hold, on the contrary, that joint action is possible in the first years of life?

To set the scene, consider that the need to be fed means that infants are involved in coordinated interactions with older caregivers from birth. In the first few months of life, infants typically appreciate dyadic interactions with others, the sort of interactions that occur when an adult spontaneously makes faces and vocalizes at an infant. This can be shown by having the infant and a caregiver interact over a video link. When instead of showing the caregiver live, a recording of her past interactions with the infant is played, infants are less satisfied (Trevarthen 1980, 323). Infants are not satisfied merely by seeing familiar gestures directed at them: even in the first months of life, they care about some aspect of the interaction.

A key milestone in development occurs in the second half of the first year, when formerly dyadic interactions expand to include 'objects, events and individuals outside of the dyad'

(Brownell 2011, 197). At this stage, infants react positively to being shown or passed objects, and you can sometimes even play at teasing them by offering and then retracting an object (compare Behne et al. 2005). But for the most part, infants in the first year of life do not initiate joint actions.

One-year-olds, by contrast, may attempt to initiate a familiar joint action by, for example, placing a cloth over their faces and waiting until an adult plays peek-a-boo (Brownell 2011). They will also attempt to re-engage adults who become distracted or unresponsive in a joint action. To demonstrate this, Warneken and Tomasello (2007, experiment 2, Trampoline) had a confederate play a game with 14-month-olds which involved bouncing a cube on a hand-held trampoline. The infants held one side of the trampoline and the confederate held the other and they bounced the cube together. After some time, the confederate froze and become unresponsive. How did the infants respond? In a minority of cases, they disengaged or attempted to continue on their own. But most of the time the infants either waited or attempted to re-engage the confederate. Their tendency to re-engage the confederate in the joint action indicates that even some 14-month-olds are to some extent aware that what they are attempting to do requires a contribution from another.1001

At around this age, children will also spontaneously initiate joint actions. Warneken, Chen and Tomasello (2006) created situations in which an adult clearly needed help. For example, the adult dropped a clothes peg and attempted, unsuccessfully to reach it. Or an adult approached and gently bumped into a cabinet with a pile of papers occupying both hands, so that she could not open it to place the papers in it. When this happened, 18-month-olds, who were positioned merely as bystanders, would spontaneously get up and provide help. They would, for example, pick up the clothes peg and hand it over, or open the cupboard for the papers to be placed. (If you access Warneken and Tomasello 2006, you can see some of the videos for yourself in the Supplementary Materials.) While Warneken, Chen and Tomasello's (2006) concern was primarily helping, many of the infants' interventions were clearly joint actions. Before they have even learnt to walk properly, humans will spontaneously initiate joint actions.

Joint action in the second year of life typically begins with extremely limited coordination on the part of children. This is neatly illustrated by (Warneken and Tomasello 2007), who created an 'Elevator' task. To obtain an object hidden in the 'Elevator', infants had to work together with an adult. One person was tasked with pushing up a cylinder from underneath a bench so that the other person could then remove the object from the cylinder above the bench. When their role was simply to take the object ('Role A'), infants were coordinated even at 14 months of age. This suggests that they can comprehend others' actions and perform complementary actions. But when the infants' role was to push up the cylinder ('Role B'), joint actions with the 14-month-olds nearly all failed and the 18-month-olds were uncoordinated with the adults. Only the 2-year-olds could push up the cylinder and hold it up while the adult retrieved the object (see Figure 15.3). Coordination of action in joint action improves substantially over the second year of life. In fact, coordination of action continues to improve gradually over several years, as Endedijk et al. (2015) demonstrate in their observations of spontaneous coordination in drumming together by 2-, 3- and 4-year-olds.

Figure 15.3 Coordination for joint action improves substantially over the second year of life

Source: Warneken and Tomasello (2007), Figure 3 (part).

We started this section by considering an inconsistent triad of claims (on p. 000). Which claim should we reject? It is clear that we cannot reject the first claim, according to which 1- and 2-year-olds are capable of performing joint actions. After all, we have just considered plenty of evidence for joint action in the second and third years of life. This leaves the second and third claims as candidates for rejection.

Some researchers suggest rejecting the third claim: they deny that a function of shared intention is to coordinate two or more agents' plans (for example, Tollefsen 2005; Pacherie 2013). This involves rejecting Bratman's account of joint action (see Section 15.3) and providing an alternative account, one in which shared intention and coordinated planning are not directly related. If we accepted one of these alternative accounts of shared intention, the claim that 1- and 2-year-olds have shared intentions would not generate the incorrect prediction that such young children are capable of coordinating their plans with those of others. This strategy—rejecting the third claim and defending an alternative account of shared intention—is the right one to adopt if you are optimistic that it is possible to give a uniform account of joint action, one that encompasses everything from joint action in infancy to the most sophisticated forms of joint action in adulthood.

Lacking such optimism, I shall pursue a different approach in what follows. I propose that we reject the second claim, according to which all joint action involves shared intention. My hope is that characterizing a form of joint action that does not involve shared intention will enable us to better understand joint actions as performed by 1- and 2-year-olds.

15.7 Collective goals vs shared intentions

Brownell neatly describes the problem we face in characterizing joint actions as performed by 1- and 2-year-olds:

All sorts of joint activity is possible without … complex reasoning … In studying its development in children the problem is how to characterize and differentiate primitive, lower levels of joint action operationally from more complex and cognitively sophisticated forms.

(Brownell 2011, 195)

The evidence we have explored so far in this chapter supports her description of the problem. One- and 2-year-olds perform joint actions (see Section 15.5). But, as we saw in Section 15.6, it seems they are several years away from coordinating their plans with those of others. This indicates that Bratman's account of joint action (see Section 15.3) does not characterize the joint actions infants perform in the second and third years of life. An alternative is needed.

What should we look for in the alternative? Minimally, it needs to distinguish the kind of joint actions 1- and 2-year-olds can perform from parallel but merely individual actions (see Section 15.1). And it would ideally be consistent with the Joint Action Conjecture (see Section 15.4).1002

As a first step, it is useful to have a distinction between collective and distributive interpretations of predicates (see Linnebo 2005, for details). Imagine Ayesha takes Zach's glass and holds it up while Beatrice pours prosecco. Unfortunately the prosecco misses the glass and soaks Zach's trousers. Here are two sentences, both true:

The tiny drops fell from the bottle.

The tiny drops soaked Zach's trousers.

The first sentence is naturally read distributively. It specifies something that each drop did individually. But the second sentence is naturally read collectively. No one drop soaked Zach's trousers; by itself, a lone drop would have no noticeable effect on his trousers at all. Rather the soaking was something that the drops accomplished collectively. If the sentence is true on this reading, the tiny drops' soaking Zach's trousers is not a matter of each drop soaking Zach's trousers.

Compare a sentence involving actions and their outcomes:

Their thoughtless actions soaked Zach's trousers.

Taken in isolation, this sentence could be read in two ways, distributively or collectively. We could imagine that we are talking about a sequence of actions done over a period of time, each of which soaked Zach's trousers. In this case, the outcome, soaking Zach's trousers, is an outcome of each of the actions. Alternatively, we could imagine several actions which have this outcome collectively. And this is the natural way of interpreting the sentence given that

we are imagining Ayesha holding a glass while Beatrice pours. In this case the outcome, soaking Zach's trousers, is not necessarily an outcome of any of the individual actions. But it is an outcome of all of them taken together.1003

Note that the distinction between distributive and collective readings involves a genuine ambiguity. To see this, ask yourself how many times Zach's trousers were soaked. On the distributive reading of the sentence just above, they were soaked at least as many times as there are actions. On the collective reading, they were not necessarily soaked more than once. So whether two or more actions involving multiple agents have an outcome distributively or collectively is not just a matter of words: the difference concerns how the actions and outcomes are related.

Consider one last sentence:

The goal of their actions was to fill Zach's glass.

Whereas the previous sentence was causal, this sentence is teleological. Whereas the previous sentence concerned an outcome which is an actual consequence of some actions, this one concerns an outcome to which actions are merely directed. Like the previous sentence, this sentence has both distributive and collective interpretations. On the distributive interpretation, each of their actions was directed to an outcome, namely, filling Zach's glass. So there were as many attempts (successful or not) to fill his glass as there were actions. On the collective reading, by contrast, it is not necessary that any of the actions considered individually was directed to this outcome. Rather the actions were collectively directed to this outcome.

Where two or more actions are collectively directed to an outcome, let us say that this outcome is a collective goal of the actions. Note two things. First, a collective goal is an outcome. It is not a mental state or a representation. Second, this definition involves no assumptions about the intentions or other mental states of the agents. This notion of a collective goal is a narrowly logical one, not a psychological one.

As Ayesha and Beatrice's pouring prosecco illustrates, some actions involving two or more people are purposive in the sense that among all their actual and possible consequences, there are outcomes to which they are directed and the actions are collectively directed to this outcome. But in virtue of what could actions involving two or more people ever have a collective goal?

Despite the simplicity of this question, there is potential for confusion. Others have used the term 'collective goal' for quite different notions. For instance, Tuomela (2002, 30–1) defines the notion in terms of 'persons believing or collectively accepting that the goal state … is to be collectively achieved'. And Gilbert (2013, 34) uses the term in connection with cases where people 'collectively espouse a certain goal, and each one is acting … in light of the fact that the goal is their collective goal'. Terminologically, these uses of the term 'collective' are unfortunate because this term is widely and consistently used in just the way I am using it in discussions about the logic of plural quantification. Substantially, the narrowly logical notion of collective goal introduced here is more basic than the other notions in this respect: if a collective goal in Tuomela's or Gilbert's or anyone else's sense were to lead to

some actions, these actions would have a collective goal in the narrowly logical sense too; but the converse is untrue. The narrowly logical notion of a collective goal allows for a clear separation between a fact that stands in need of explanation (that some actions are collectively directed to an outcome) and the things which putatively explain it (such as collective espousal, according to Gilbert).

So in virtue of what could actions involving two or more people ever have a collective goal? To illustrate, Ayesha might say, truthfully, ‘The collective goal of our actions was not to soak Zach’s trousers in sparkling wine but only to fill this glass.’ What could make Ayesha’s statement true?

One way to answer the question is by invoking a notion of shared intention. Suppose Ayesha and Beatrice have a shared intention that they fill the glass. Then, on many accounts of shared intention including Bratman’s (see Section 15.3), the shared intention involves each of them intending that they, Ayesha and Beatrice, fill the glass; or each of them being in some other state which picks out this outcome. The shared intention also provides for the coordination of their actions (so that, for example, Beatrice doesn’t start pouring until Ayesha is holding the glass under the bottle). And coordination of this type would normally facilitate occurrences of the type of outcome intended. In this way, invoking a notion of shared intention potentially explains what it is in virtue of which actions involving two or more people could have a collective goal.

But invoking shared intention is exactly what we want to avoid (see Section 15.6). Are there also ways of answering the question which involve psychological structures other than shared intention?

One possibility is to invoke expectations about collective goals.

15.8 Expectations about collective goals

Suppose that Ayesha and Beatrice each reasonably expect the actions they are about to perform to have the collective goal of filling Zach’s glass, and that they act on these expectations. Then their expectations specify a collective goal and provide for the coordination of their actions in much the way that a shared intention would. After all, someone who acts on an expectation that her actions and those of another will have a collective goal thereby has an incentive to coordinate her actions with the other’s insofar as this will facilitate the occurrence of the collective goal. Apparently, then, an outcome could be a collective goal of Ayesha and Beatrice’s actions in virtue of them each acting on the expectation that this outcome will be a collective goal of their actions.

This suggests a minimally demanding sufficient condition for joint action: ‘Where an event comprises two or more agents’ actions and the actions have a collective goal in virtue of the agents’ acting on expectations that these actions will have this collective goal, the event is a joint action.’ This sufficient condition enables us to distinguish paradigm contrasts between joint action and parallel but merely individual actions (see Section 15.1). For instance, earlier we contrasted two friends out walking together in the way friends typically walk together with two strangers who happen to be walking side by side. The friends but not the strangers meet

the above sufficient condition for joint action: they are plausibly acting on an expectation that their walking together will be a collective goal of their actions. We also contrasted several performers simultaneously running to a central shelter in order to perform a dance with an event involving park visitors running to the central shelter in order to escape a storm. Again, the strangers do not meet the above sufficient condition whereas the performers plausibly do—plausibly the performers are each running to the shelter on the expectation that this action together and the others’ actions have the collective goal of performing the dance.1004

Could 1- and 2-year-olds meet the proposed sufficient condition for joint action? A collective goal is merely an outcome to which some actions are directed. The goal-tracking abilities infants manifest in the first year of life show that they can form expectations about actions being directed to goals (see Chapter 10). It is only a small step to the further conjecture that infants can track goals to which two or more agents’ actions are collectively directed. And there is even some evidence in support of this conjecture (Fawcett and Liszkowski 2012). So although we cannot be certain, it is reasonable to assume that 1- and 2-year-olds can form and act on expectations about collective goals. And this is all meeting the proposed sufficient condition for joint action requires.

It may be objected that the proposed sufficient condition for joint action is not actually sufficient at all and that an additional requirement is needed, namely, that the agents have common knowledge concerning their expectations. This might make the claim that 1- and 2-year-olds can meet the sufficient condition for joint action less plausible. It is one thing to have an expectation concerning a collective goal of some actions and quite another to know that someone has such an expectation. Further, if infants’ expectations concerning collective goals involve motor and perceptual representations rather than knowledge states (see Chapter 11), it might turn out that they are not in a position to have knowledge concerning these expectations. It is therefore worth asking why the requirement on common knowledge needs to be added to the proposed sufficient condition for joint action. Philosophers who appeal to common knowledge in characterizing joint action have been strikingly reticent about why they do so.1005 And some have argued that common knowledge is not in fact necessary at all (see Blomberg 2016). As things stand, then, there does not appear to be a compelling argument to show that common knowledge is needed in giving sufficient conditions for joint action. Rather than being an essential feature of joint action, various kinds of common knowledge may be optional coordination enhancers.

A similar line of thought applies to the claim that commitment is an essential characteristic of all joint action (Gilbert 2013). Children appear to first become sensitive to commitments in the context of joint action at around 3 years of age (Gräfenhain et al. 2009). So adding requirements on commitments to a sufficient condition for joint action may well have the consequence that children in the first years of life cannot meet the sufficient condition. But commitment, like common knowledge, is plausibly an optional enhancement rather than an essential feature of all joint action (compare Bratman 2014, 118–20).

Overall, then, it is not impossible that the proposed sufficient condition for joint action will turn out to be actually sufficient. At least we do not currently have reason to assume that it needs supplementing by an appeal to common knowledge or commitment, thereby rendering it unsuitable for understanding 1- and 2-year-olds’ joint actions. Let us therefore assume that

the proposed sufficient condition for joint action is actually sufficient for joint action, and that it is one that 1- and 2-year-olds could meet. Now consider a further question. Does the proposed sufficient condition characterize joint actions as performed by children in the first years of life?

One point in favour of the proposed sufficient condition is that it appears to be consistent with the evidence on joint action in the first years of life considered in Section 15.6. As we saw, 1-year-olds can spontaneously initiate joint actions. If we think of joint actions as characterized by expectations about collective goals, imitating a joint action is a matter of doing things which create such expectations. And children may initially do things which create expectations about collective goals in themselves and in others without necessarily being aware of doing so. One-year-olds also attempt to re-engage others in a joint action following an interruption. This is consistent with their having expectations about collective goals. A child with an expectation that her own and another’s actions will have a collective goal may also realize that the collective goal is unlikely to be realized without the other’s actions, which would give her an incentive to re-engage the other. Overall, then, the evidence currently available on joint action in the first years of life is consistent with the claim that such joint actions are characterized by expectations about collective goals.

This argument should give us little confidence that joint actions as performed by children in the first years of life really can be characterized simply by expectations about collective goals. The recent history of developmental psychology is all about how using more sensitive measures such as anticipatory looking or pupil dilation reveals that infant cognition is unexpectedly sophisticated (see Chapter 10, for examples). A more cautious claim would be that joint action in the first years of life is partially characterized by expectations about collective goals. This claim is at least a reasonable starting point in attempting, over the remaining chapters, to understand how joint action may play a role in explaining the emergence of knowledge in development.

15.9 Conclusion

Our aim in this chapter was to understand joint actions as performed by children in the first years of life. Such joint actions are a matter of two or more agents each acting on expectations that all of their actions will have a certain collective goal. This characterization of the kind of joint action 1- and 2-year-olds perform provides some of the groundwork for investigating the possibility that, as the Joint Action Conjecture states, abilities to perform joint actions play a role in explaining the developmental emergence of knowledge.

How did we arrive at the idea that joint action in the first years of life is characterized by expectations about collective goals? We started by considering Bratman’s account of joint action (in Section 15.3). This is the best developed and most influential account, and several developmental scientists have argued that it correctly characterizes joint actions as performed by children in the first years of life. But if they are right, we must reject the Joint Action Conjecture. For, as we saw (in Section 15.4), the following claims form an inconsistent triad:

1 Abilities to perform joint actions play a role in explaining the developmental emergence of knowledge, including knowledge of others’ minds. (This is the Joint Action Conjecture.) 2 All joint action involves shared intention. 3 A function of shared intention is to coordinate two or more agents’ plans (as Bratman’s account implies).

The inconsistency arises because abilities to coordinate plans with another presuppose abilities to know things about others’ minds (see Section 15.4). Claims nos 2 and 3 therefore entail that invoking abilities to perform joint actions would be to presuppose that infants already have knowledge of minds, forcing us to reject claim no. 1.

Recognizing that the triad of claims is inconsistent led us to ask which claim should be rejected. The view that Bratman’s account of joint action characterizes joint actions as performed by children in the first years of life leads to the prediction that such children should be capable of coordinating their plans with those of others. But this prediction is incorrect: the first signs of coordinated planning emerge much later in development, at around 5 years of age (see Section 15.5). We must therefore reject the view that joint action in the first years of life is characterized by Bratman’s account of shared intention. Consequently, if we maintained both claims nos 2 and 3 of the inconsistent triad just above, then we would have to deny that children in the second and third year of life are capable of joint action.

But this we should probably not do. A variety of evidence indicates that although they have quite limited capacities to coordinate their actions with others, even 14-month-olds will spontaneously initiate joint action with an adult. Children of around this age also demonstrate awareness in the context of joint action that success requires another person’s contribution (see Section 15.6). On balance, then, it does not seem plausible to deny that children in the second and third year of life are capable of joint action. We must therefore reject at least one of claims nos 2 and 3 of the inconsistent triad. Doing so requires us to find an alternative to the view that Bratman’s account characterizes joint action in the first years of life, one consistent with children’s capacities and their limits.

To this end I introduced the notion of a collective goal. A collective goal is any outcome to which the actions of two or more agents are collectively directed. And for their actions to be collectively directed to an outcome is for their actions to be directed to an outcome where this is not, or not only, a matter of each action being individually directed to that outcome. Unlike the deeply controversial and hard-to-ground notion of shared intention, the notion of a collective goal is a narrowly logical one (see Section 15.7). The step from the logical to the psychological occurs when we ask in virtue of what some agents’ actions could have a collective goal. One possibility is that the agents each act on an expectation that all of their actions will have a collective goal. This suggested a minimally demanding sufficient condition for joint action involving agents acting on expectations about collective goals (see Section 15.8). This sufficient condition in turn provides a candidate characterization of joint action in the first years of life, one capable of distinguishing paradigmatic contrasts between joint actions and parallel but merely individual actions. Perhaps 1- and 2-year-olds’ joint actions are actions they perform in acting on expectations about collective goals.

If this is right, we should question not only Carpenter’s (2009) view that joint action in the first years of life is correctly characterized by Bratman’s account of shared intention but also broader appeals to shared intentionality. Tomasello and Carpenter (2007, 124) propose that a capacity for ‘shared intentionality’ emerges ‘at around the first birthday’ from ‘a uniquely human line of development for sharing psychological states with others’. Reflection on 1- and 2-year-olds’ joint actions suggests that we may not need theoretically disputed and hard-to-pin-down notions of sharing in order to characterize what they are doing. Perhaps it is not shared intentionality, whatever exactly that turns out to be, but a set of minimally complex abilities, including abilities to form and act on expectations about collective goals, that matters for understanding cognitive development from the second and third years of life.

This characterization of joint action, unlike Bratman’s, is consistent with the Joint Action Conjecture. This is because Bratman’s requires coordinating plans, which requires knowledge of other’s mental states. By contrast, characterizing joint action by appeal to collective goals places no demands on knowledge and so opens the way to explaining the developmental emergence of knowledge, including knowledge of others’ minds by appeal to 1- and 2-year-olds’ abilities to perform joint actions.

Notes

16 Conclusion to Part II

Abilities to track actions and mental states, and to perform joint actions, appear about as early in development as abilities concerning physical objects. This should be no surprise given how social most humans are. The real headline from the discoveries reviewed in Part II is about limits, not successes.

These limits—patterns in what infants can and cannot do—mean that when attempting to characterize infants’ abilities in each domain, we are immediately confronted by developmental puzzles. One kind of puzzle concerns abilities whose manifestation is response-dependent. For goal tracking and mindreading, there are single scenarios in which infants will manifest a goal-tracking or mindreading ability when one response is measured but not when another response is selected. For example, contrast pupil dilation with anticipatory looking in the case of goal tracking (Gredebäck and Melinder 2010), or anticipatory looking with an explicit question in the case of mindreading (for example, Low et al. 2014). One puzzle facing any theory is to explain why such systematic discrepancies in performance should occur.

By itself, this is a shallow kind of puzzle. After all, mere response-dependence may invite a methodological explanation. A deeper puzzle arises because patterns of response-dependence interact with other limits. Thus, goal tracking is sometimes, but not always, limited by an infant’s own abilities to perform actions; and within this own-action limit, response-dependence does not apply, or does not apply in the same way (see Section 10.7). Equally, mindreading is sometimes, but not always, subject to a signature limit concerning numerical identity (see Section 14.7); and whether the limit is observed depends on which kind of response is measured (see Section 14.8). One challenge for any developmental theory of goal-tracking or mindreading is to explain, ideally in a way that generates readily testable novel predictions, why there are these interacting limits on abilities.

Whereas puzzles about goal tracking and mindreading arise prior to accepting any theory, the main puzzle about joint action was linked to a particular theory. On Bratman’s theory, all

joint action involves shared intention and having shared intentions implies being able to coordinate planning, which in turn implies being able to know things about another’s mental states (see Section 15.3). The puzzle about joint action arises because 1- and 2-year-olds appear capable of joint action, but incapable of coordinating planning (as we saw in Section 15.5).

We should have low confidence that the solutions to puzzles about actions, minds and joint actions we have considered will turn out to be correct. There is little prospect of firmly establishing a developmental theory on narrowly theoretical grounds, or by retrofitting it to past discoveries. Ultimately everything turns on whether the theory’s predictions turn out to be broadly correct. Nevertheless, some aspects of the theoretical approach could turn out to have lasting value even if many details turn out to be wrong. It is therefore worth attempting to draw together what we have learnt about goal tracking, mindreading and joint action.

16.1 Dual process theories

The general approach to solving the puzzles about goal tracking and mindreading we have followed leans on the idea of a dual process theory. The core claim of any dual process theory is just that humans’ abilities in a domain involve two or more processes; and the processes are distinct in this sense: the conditions which influence whether one process occurs differ from the conditions which influence whether another occurs.

The term ‘dual process’ can be misleading. It may be understood as meaning there are exactly two kinds of process, whereas we have often distinguished more than two. I think the term remains apt, though, because the difficult thing is to get people past one. Once you get past the idea that goal tracking, mindreading and anything else we pre-theoretically think about as a single ability has to be explained by appeal to a single kind of process, the step from two processes to three or more is not so hard.

The most important duality we have observed is linked to knowledge. On the one hand, there are inferential processes involving knowledge states. On the other hand, there are distinct motor and perceptual processes, and early-developing belief-tracking processes (whatever exactly those are). In every domain, the inferential processes emerge later in development, months or even years after the early-developing processes are manifest. This makes it theoretically coherent to appeal to the early-developing motor and perceptual processes in explaining the developmental emergence of knowledge.

There is an important difference in the dual process theories we have considered. Some, but not all, involve a further distinction between two or more kinds of early-developing processes. In the domain of action, we could distinguish two kinds of process which seem to occur even in the first year of life. We have evidence for proper goal tracking based on motor processes (see Section 11.3). And we have evidence that infants can also track targets thanks to whatever broadly perceptual processes underpin perceptual animacy. Similarly, in the domain of physical objects, we distinguished the operations of a system of object indexes from motor processes (see Chapter 6): both are present in the first months of life. By contrast, in the domain of minds we have yet to see evidence that distinct belief-tracking

processes occur even in the first year of life. Perhaps this is because less is currently known about the nature of these processes, or perhaps it is an artefact of our narrow focus on tracking beliefs rather than mental states generally. Or maybe mindreading is just more unified than abilities in other domains.

16.2 Pluralism about models

Alongside dual process theories, another key feature of our approach is pluralism about models. A model is just a way some part or aspect of the world could be.

There is a striking contrast between developmental research on physical objects and developmental research on minds and actions.

In the case of the physical, it is hardly questioned that processes for segmenting and tracking objects rely on distinctive models of the physical, models which are in some respects less accurate. Indeed, it was unnecessary to discuss models explicitly in Part I. It is too obvious to need emphasis that Spelke’s Principles of Object Perception characterize a distinctive model of the physical, one that is less accurate but also less complex that those which adult humans typically rely on in ordinary thought and talk about physical objects. And of course the models of the physical relied on in everyday life are in turn less accurate than models deployed by engineers or physicists. The idea that distinct processes underlying abilities concerning physical objects could involve distinct models of the physical is hardly controversial and has attracted scarce attention (but see Kozhevnikov and Hegarty 2001 for a particularly elegant exception).

In the case of the mental, I often encounter resistance to the idea that distinct mindreading processes might rely on different models of minds and actions. It may be tempting to assume that there is just one model of the mental and all mindreading involves the use of that model. Or, more carefully (to accommodate insights about development such as those by Wellman, Fang and Peterson 2011), the tempting assumption is that all models of the mental comprise a family in which one of the models, the best and most sophisticated model, contains everything contained in any of the models.

I suspect this assumption is tempting for at least two reasons. One possible reason is lack of familiarity with alternative models of minds and actions. Even the slightest familiarity with the history of science enables us to get a fix on how impetus-based and Newtonian models of the physical differ, for example. There is no comparably easy way to get a fix on the different models of minds and actions (but see Section 14.6).

The other possible reason it may be tempting to assume that there is just one model of minds and actions arises from self-understanding. It may be quite difficult to accept that the model of minds and actions which underpins your everyday thinking about yourself and those around you is not especially accurate. But note that this model of minds and actions has to serve a variety of purposes. As well as enabling you to predict, it underpins normative, ethical and legal activities.1006 You might judge a particular line of thought to be a blameless mistake; or you might blame someone for an emotional response which, although inappropriate, is not in any sense a mistake. And, of course, legal questions can hinge on matters of

belief and intention. Insofar as a model of minds and actions serves all these purposes, should we expect it to be accurate? My guess is that a model’s effectiveness in allowing us to knowingly reach normative, ethical and legal verdicts will often be in tension with its accuracy, given the limits on our cognitive resources. But even if a model of minds and actions were only used for prediction, the need to generate timely predictions requires speed–accuracy trade-offs, and these would likely mean that the model trades accuracy for simplicity (Section 14.4).

The struggle to get from Aristotelian to Newtonian models of the physical—to get beyond the idea that informal observation and guesswork provide the only way to characterize physical phenomena—was notoriously long and bitter. Cleaving to the idea that whatever model underpins adults’ most reflective thinking is the one and only accurate model of minds and actions is a relic of this struggle. To understand development, we need to distinguish not only processes but also the models these processes rely on. By characterizing a model relied on by early-developing mindreading or goal-tracking processes, we can see how minds and actions appear from the perspective of an infant.

16.3 Goal tracking is the foundation

Tracking the goals of actions is indispensable for mindreading. After all, it is predictions and retrodictions of action that provide the evidential basis for ascribing mental states—this holds even for the states involved in minimal theory of mind (see Section 14.6). And goal tracking is, of course, essential whenever joint action (or any form of social interaction) involves monitoring another’s actions.

The foundation for research on goal tracking is provided by Gergely et al.’s (1995) Teleological Stance which states, roughly, that goals are those outcomes which an action is best suited to bringing about. This formulation can be extended by adding further to principles that implicitly specify a relation between actions to outcomes. Where the principles obtain, then the outcomes will, often enough, be goals of the actions.

Given that the Teleological Stance is formally adequate, we can be sure that pure goal tracking is possible in principle. That is, goal tracking can occur independently of any information about mental states. Theoretically, this matters because it shows that a theory of goal tracking can be independent of, and therefore foundational for, theories of mindreading and of joint action. Practically, focusing on pure goal tracking allows us to formulate and evaluate conjectures about goal tracking independently of conjectures about mindreading.

Adopting the Teleological Stance leaves open the question, Which processes enable infants to track goals? (This is a counterpart of the Linking Problem encountered in Chapter 4; see Section 11.3.) Reflection on the limits of goal tracking in adults and infants provides indirect support for the view that not all goal tracking involves inferential processes operating on knowledge states: instead, at least some goal tracking involves motor processes only (see Section 11.2). One problem for this view is that it is incompatible with many cases in which infants in the first year of life appear to adopt the Teleological Stance but could not be relying on motor processes. But it is possible that in these cases the infants are merely tracking

targets and relying on heuristics rather than adopting the Teleological Stance. If so, their abilities in these cases may be explained by whatever underpins perceptual animacy in adults (see Section 11.4).

In this Part, we have travelled from goal tracking to joint action. This may give the impression, wrongly, that a complete theory of goal tracking could be independent of theories of joint action or mindreading. Although goal tracking is a foundation for mindreading and joint action, it is also possible that fully understanding goal tracking may require us to travel in the opposite direction. We may also need to understand how some forms of goal tracking depend on joint action.

16.4 When joint action enables goal tracking

The Teleological Stance is limited. Using the Teleological Stance hinges on you (or motor processes in you) computing means-ends relations. But suppose that you are ignorant of the relevant means-ends relations. Perhaps, for instance, you are observing someone using a novel tool and you have no idea what it is for or how it might work. Or maybe the action you are observing involves multiple steps that do not form a familiar sequence, can occur in various orders and can be interspersed among other activities. In such situations you cannot use the Teleological Stance because the means are opaque to you.

You will also run into the problem of opaque means if you have no prior experience of referential communication. Communicative actions characteristically have goals which the actions are the means to realizing only because others recognize them as the means to realizing those goals (that is, they involve a Gricean circle). This makes the means-ends relation involved to a non-communicator, posing a significant problem for attempts to explain the developmental emergence of communicative abilities.

Just here joint action is relevant. Abilities to perform joint actions provide you with a way to track the goals of other’s actions, which does depend on having information about the relevant means-ends relations. Joint action can therefore enable you to overcome the problem of opaque means.

How? According to Butterfill (2012a), the following inference characterizes a route to knowledge of others’ goals:

1 You are willing to engage in some joint action or other with me. 2 I am not about to change the single goal to which my actions will be directed.

Therefore:

3 A goal of your actions will be my goal, the goal I now envisage that my actions will be directed to.

Call this inference your-goal-is-my-goal. To say that it characterizes a route to knowledge implies two things. First, in some cases it is possible to know the premises, 1–2, without

already knowing the conclusion, 3. Second, in some of those cases, knowing the premises would put one in a position to know the conclusion.

Consider each point in turn. Because another can signal that they are about to engage in joint action with you by, for example, smiling and making eye contact, it is possible to know that they are about to engage in joint action with you independently of any information about the goals of the joint action. And because your goals are often enough obvious to people around you, and they are often enough willing to help you achieve them, the step from premises to conclusion is sometimes reasonable. (Note that there is no guarantee: the inference is not supposed to be deductive, of course.)

If the above inference really is a route to knowledge about others’ goals, then it shows how abilities to perform joint actions enable a kind of goal tracking that can avoid the problem of opaque means. The idea is not, of course, that infants (or adults) who use this kind of goal tracking really need to know premises nos 1 and 2. No knowledge is needed at all, just as the Teleological Stance characterizes a form of goal tracking which can be implemented with motor representations and no knowledge at all (see Section 11.2).

The theory of goal tracking has been treated as an input to theories of mindreading and joint action in this Part. This simplification makes sense, given how fundamental goal tracking is. But, as the problem of opaque means and the possibility of avoiding it by exploiting abilities to perform joint actions suggest, we will eventually also need to consider how 1- and 2-year-olds’ abilities to perform joint actions and to track mental states can facilitate goal tracking.

16.5 Joint action and the developmental emergence of knowledge

We have seen that infants in the first year of life can track actions and mental states and perform joint actions. None of these abilities involve knowing facts about particular actions, mental states or intentions; nor do they require planning abilities. It is therefore safe to work on the premise that these abilities, far from presupposing that knowledge is already in place, are somehow implicated in the emergence in development of knowledge. Knowledge of particular minds and actions, like knowledge of particular physical objects, is not manifest in the very first months and years of life but somehow emerges from core knowledge and abilities for joint action.

Note

17 Conclusion

I started with a simple question. How do humans first come to know about objects, actions and minds? On the view elaborated in this book, the answer, whatever it ultimately turns out to be, will involve two primary ingredients— core knowledge and joint action—where the role of core knowledge is to make joint action possible.

This view comes out of an exploration of what has been discovered about the developmental emergence of knowledge in two domains: physical objects (in Part I) and minds and actions (in Part II). My conclusion is that infants’ surprisingly sophisticated abilities concerning objects, minds and actions do not involve knowledge at all. Instead they involve a combination of broadly perceptual, motoric and metacognitive states and processes, perhaps among other things. These states and processes do not appear to change over development: they are also found in adults, and they underpin much the same abilities in adults as in infants. The fact that these states are distinct from knowledge enables us to invoke them without circularity in explaining the developmental emergence of knowledge. The early-developing abilities must somehow provide a basis for the later acquisition of knowledge. But how? Any attempt to answer this question must acknowledge a problem. The states which underpin early-developing abilities are cut off from knowledge: in both adults and infants, they exhibit intentional isolation from, and lack of inferential integration with, knowledge. These states influence the rest of the mind only indirectly, by their influences on metacognitive feelings, behaviours and other intentional isolators. There is no prospect, therefore, of postulating operations on the contents of these states which will transform them into the contents of knowledge states. There can be no direct representational connections between the states underpinning early-developing abilities and knowledge states. The emergence of knowledge has to be a process of rediscovery, of discovering as if for the first time things about objects, minds and actions which in some sense infants have long relied on in their interactions with the world. While we know almost nothing about how rediscovery occurs, it is probably not

something infants achieve alone. Rediscovery must be a joint action. Just here we face an objection: on the leading, best-developed views, joint action presupposes knowledge (and much further sophistication besides) and so could not explain its emergence in development. One candidate solution is to reject the leading views in favour of an account of joint action in terms of collective goals. Thinking in terms of collective goals allows us to see how infants’ early-developing abilities to track aspects of minds and actions provide a basis for the social skills that make rediscovery possible.

The developmental emergence of knowledge depends on non-social, early-developing abilities concerning objects, minds and actions; so cannot be understood in entirely social terms. But nor can it be understood by ignoring infants’ interactions with others entirely, for the states underpinning those early-developing abilities are intentionally and inferentially isolated from knowledge. This creates a special, as yet barely explored role for joint action in development. Coming to know your first facts about particular objects, minds and actions is a process of rediscovery, which is a joint action.

These conclusions are controversial and rest on some quite wild conjectures whose predictions are yet to be tested. But even if all of them were accepted, we would still be some way from answering the question about the developmental emergence of knowledge. It is therefore worth highlighting some themes running through the discoveries and puzzles we have considered, with a view to how these might eventually inform constructing an answer to the question about how humans first come to know simple facts about particular objects, actions and minds.

17.1 Infants rely on minimal models ...

A model is a way some part or aspect of the world could be. Models can be used not only to characterize the world, but also to characterize how the world appears from the point of view of a particular individual (or processes within her). A significant part of understanding the developing mind involves constructing models. Spelke’s Principles of Object Perception characterize a model of the physical (see Chapter 2), and Gergeley and Csibra’s Teleological Stance characterizes a model of action (see Chapter 10); similarly, minimal theory of mind identifies a model of minds and actions (see Chapter 14). Insofar as these models are descriptively adequate, they enable us to see objects, minds and actions from an infant’s point of view.

One striking feature of all the models we have considered is that they are unsophisticated, being characterized by a few principles which are at best rough approximations to the truth. These are minimal in the sense that no less sophisticated models have been described.

Unsophisticated and inaccurate models are often useful in a limited range of circumstances. In the case of the physical, an impetus model of the physical works as well as any other as long as you remain on the surface of this planet and avoid launching things vertically; and a Newtonian model is fine if you are not travelling too fast. Less sophisticated models often have an advantage for finite minds under time constraints: their use demands less cognitive power. A minimal model can make possible more extreme speed–accuracy trade-offs that would otherwise not be available.

Compared to adults, infants have limited memory, attention and inhibitory control. They also lack some of the external supports to thinking that adults benefit from: they cannot talk to themselves, write notes or make diagrams, for instance. Further, their experience of the world is more circumscribed and they have had fewer opportunities to automatize patterns of acting and thinking. These are all reasons why an infant might ideally be able to make a different speed–accuracy trade-off than would a reflective adult who is not under any time pressure. This explains why we find that models involved in infants’ early-developing abilities concerning objects, minds and actions are minimal.

17.2 ... As do adults, sometimes

Development does not seem to be a process of elaborating these models, turning the minimal into the more sophisticated. In considering objects (Section 6.1), actions (Section 11.2) and minds (Section 14.7), we found evidence that some processes in adults rely on the very minimal models that underpin infant’s abilities.

These findings can also be explained by appeal to the role of models in enabling speed–accuracy trade-offs. If minimal models merely became more elaborate throughout development, this would restrict the range of speed–accuracy trade-offs available to older children and adults. Retaining the minimal models alongside more sophisticated models enables greater cognitive flexibility.

Just identifying a model counts as progress in understanding developing minds because it enables us to see things from an infant’s (and a cognitively constrained adult’s) point of view. But any time we identify a model, we are confronted with a challenge. What links the model (or the principles we use to characterize the model) to the mind of an infant? Answering this question is a challenge because of the puzzles developmental discoveries consistently throw up.

17.3 Puzzles matter

In tracing many discoveries about development, we have consistently been confronted by puzzles. Where a scientist might view puzzles as a sign of failure, a philosopher is likely to see identifying and elucidating puzzles as an end in itself.

Any attempt to answer questions about when humans first know simple facts about particular objects, actions or minds by considering the existing discoveries is quickly confronted by a puzzle of this form: some observations indicate that knowledge is in place in the first months of life and that there is no developmental transition between infancy and later years; whereas other observations indicate the contrary, and provide evidence for a developmental transition; further, the difference between the observations arises from apparently irrelevant factors such as a difference in response type measured (from pupil dilation to anticipatory looking, for example), timing constraints imposed, or scenario used (occlusion to endarkening, or location to numerical identity, for example).

In each of the domains we considered—objects, minds and actions—there is a clear temptation to dismiss the puzzle by insisting that one or the other side involves extraneous methodological defects. But attempts to identify these extraneous methodological defects have not been successful, and each side can point to a variety of impressive findings.

We should resist this temptation. The findings appear puzzling only as long as we assume, perhaps tacitly, that whatever states underpin the abilities being measured are like knowledge in that they must, in the absence of an impediment, be capable of guiding any response. But this assumption is neither theoretically mandatory nor supported by much (if any) evidence. The alternative, as we have seen, is to treat the puzzles as informative both about the nature of the states and processes involved and also about the models that underpin these abilities. The apparent puzzles, properly understood, may point to a signature limit of motor representation (which is unperturbed by brief endarkening but strongly affected when occlusion events involve impenetrable barriers; Section 6.7), or of minimal models of mind (Section 14.7), or of perceptual animacy (which may enable pupil dilation but not anticipatory looking; Section 11.4), among other things.

The puzzles are so valuable because mere successes taken in isolation could be explained in indefinitely many ways, while an isolated observation of mere failure can be hard to distinguish from a null result. The puzzles provide insights about limits on success. These limits are what allow us to falsify hypotheses about which models, states or processes underpin the successes. This is the key to solving linking problems.

17.4 Linking problems abound

One theme running across all domains we have considered is that infants demonstrate surprisingly sophisticated abilities surprisingly early in life. These enable them to manifest symptoms associated with knowledge, which makes it tempting to conjecture (as many have) that infants’ abilities are based on knowledge or belief. Such conjectures are versions of what I called the Simple View.

One attraction of a Simple View is that it provides a straightforward answer to the problem about what links a model to the mind of an infant: the principles which characterize the model are things the infant knows.

No Simple View is to be discarded lightly because there is invariably much going for it. In each domain we considered, the Simple View required no theoretical novelty, generated readily testable predictions and was supported by some evidence. But versions of the Simple View do quite consistently generate incorrect predictions (as we saw in Chapters 4, 10 and 12). Infants manifest symptoms associated with knowledge. But whatever underpins these manifestations cannot be knowledge because there are limits on which responses it can guide or because it lacks the inferential integration characteristic of knowledge (see Section 1.2).

At this point we are confronted with a linking problem (see Chapter 4). We have identified some principles that are descriptively adequate to infants’ early-developing abilities in a domain. If not knowledge or belief, what does link these principles to their minds?

As background to this question, we have been working with a crude picture of the mind (introduced in Section 1.3). The mind comprises at least three kinds of states and processes:

1 epistemic (that is, knowledge-related); 2 motoric; 3 perceptual.

Although crude, this picture of the mind has the virtue that all three kinds of state postulated are needed by large bodies of theoretical work. If we seek to go beyond it by adding further kinds of state, we need to be sure that these novel kinds of state are characterized in ways that are both theoretically coherent and empirically motivated.

The failure of the Simple View invites us to go beyond the crude picture of the mind. After all, infants’ abilities did not initially seem to be entirely a consequence of perceptual or motor representations. So we can easily get the impression that the crude picture is insufficient for us to solve linking problems. Yet existing attempts to go beyond the crude picture have been challenged on theoretical grounds, and they do not seem to generate novel predictions (as we saw in Chapter 4). That is a large part of why linking problems are so difficult: they seem to require a way of going beyond the crude picture.

17.5 Core knowledge isn’t what you think it is

One attempt to solve linking problems is associated with the notion of core knowledge. This is often understood as a fourth kind of mental state, something distinct from knowledge proper as well as from perceptual and motor states. In this case, postulating core knowledge means adding a novel kind of mental state, a ‘third [fourth] type of conceptual structure’ (Carey 2009, 10).

I have argued against such a construal on the grounds that existing theoretical accounts of core knowledge do not pull their weight by generating novel predictions (Section 5.2); nor are things improved by appealing to standard theories of modularity. None of these arguments are conclusive, of course. But they do motivate seeking alternatives. And, as we saw in Section 8.2, there is at least one alternative way of characterizing core knowledge. This is a lighter (theoretically modest) approach which appears to require fewer bold commitments while being no less useful for tackling the puzzles.

On the lighter approach, ‘core knowledge’ is initially just a label for whatever it is that underpins early-developing abilities in a domain. Issues like whether core knowledge is innate, encapsulated or unchanging across development are not treated as a matter of stipulation but left over for discovery. We make progress by identifying core knowledge with states whose existence could be independently established.

On the lighter approach, core knowledge is not a solution to any linking problem. It is a label for a solution yet to be identified (and perhaps a guess at the rough shape the solution will take).

None of this means that theories of core knowledge have nothing to offer. Innateness aside (for reasons explored in Chapter 9), many of the insights which prompted introducing core knowledge appear to be robust. One key insight is that what underpins infants’ abilities also plays a role in cognition throughout life and thus can be found in adults (Spelke 1994; Carey and Spelke 1996). Another key insight is that core knowledge is not inferentially integrated with knowledge proper. Perhaps these insights do not hold good across every domain in which infants manifest abilities, but they do seem to characterize those domains we have considered. Accepting the lighter approach to core knowledge does not entail rejecting these insights: the point is just that the insights do not by themselves require or enable us to go beyond the crude picture of the mind by postulating novel kinds of mental states or processes. The puzzles and the linking problems remain unresolved.

17.6 How to solve linking problems

The inspiration for the approach to solving linking problems pursued here is the CLSTX Conjecture (from Section 6.3). According to this conjecture, what links the Principles of Object Perception to infants’ (and adults’) minds is a system of object indexes. This conjecture does not require postulating novel kinds of mental states or processes. The postulation of object indexes is independently motivated by discoveries and theories in non-developmental areas of cognitive science, and they are relatively well understood.

The success of the CLSTX conjecture suggests an approach to solving linking problems: go beyond the crude picture only when there is independent reason to postulate a novel kind of state or process. In line with this approach, we have considered conjectures invoking object indexes, motor representations of objects’ affordances, metacognitive feelings, motor representations of outcomes in action observation and perceptual animacy. Although none of these is strictly part of the crude picture of the mind, their postulation is independently motivated, they are relatively well understood and recognizing their existence involves no radical departure from the crude picture.

The CLSTX Conjecture turned out to generate incorrect predictions, of course (Section 6.6). But the approach it inspired continued to work: successive revisions introduced motor representations concerning affordances (Conjecture O from Section 6.8) and metacognitive feelings (Conjecture Oᵐ from Section 7.6). Similarly, it was possible to explain how principles about actions are linked to infant’s minds by appeal to a combination of motor representations and perceptual animacy.

We do not (yet) need to postulate novel kinds of representation or cognitive process in order to solve linking problems.

Whereas proponents of core knowledge standardly construe it as a novel ‘type of conceptual structure’ (Carey 2009, 10), the evidence so far suggests that we can decompose core knowledge into perceptual and motoric constituants plus metacognitive feelings that are already familiar from research in non-developmental cognitive science (see Section 1.3).1007 This implies that core knowledge lacks unity. Rather than labelling one thing, core knowledge of objects turns out to comprise a variety of things including, among others, object indexes and

motor representations. Core knowledge also turns out to lack uniformity: the states and processes it labels in one domain are not the states and processes it labels in every other domain. This lack of unity and uniformity is part of what makes theorizing about development so hard. What seems to be a single ability rarely involves just one kind of state or process (there is a lack of unity), and what goes for one domain rarely goes for another (there is lack of uniformity).

As mentioned back in Chapters 1 and 4, Davidson expresses scepticism about the possibility of solving linking problems:

If you want to describe what is going on in the head of the child when it has a few words which it utters in appropriate situations, you will fail for lack of the right sort of words of your own. We have many vocabularies for describing nature when we regard it as mindless, and we have a mentalistic vocabulary for describing thought and intentional action; what we lack is a way of describing what is in between.

(Davidson 2001, 127–8; my emphasis)

I think Davidson is wrong, although I appreciate the encouraging prognosis.1008 Combining minimal models with insights into perceptual, motor and metacognitive processes enables us to describe even what is going on in the heads of wordless infants.

17.7 Representation: handle with care

The conjectures we have considered all involve representation in one way or another. As we saw in Section 4.5, Haith (1998) claims that ‘no concept causes more problems in discussions of infant cognition than that of representation’. I disagree; or at least I don’t think it is at all necessary for the concept of representation to cause problems. We can avoid problems by bearing in mind three uncomplicated theoretical points.

First, representing is distinct from tracking. Not all tracking involves representing (Section 13.1). What is more directly observed in an experiment is usually tracking rather than representing. Supporting conclusions about representation using observations concerning tracking is difficult but not impossible. To make such an inference, your hypothesis about representation has to generate novel predictions. That is, it has to add to the theory rather than merely providing a way of restating a claim about tracking. The method of signature limits (from Section 6.4) provides one way to get from tracking to representing.

Second, hypotheses about representation should avoid, whenever possible, postulating representation without postulating a particular kind of representation (see Section 4.5). Representations are always perceptual, motoric, epistemic, photographic, cartographic, or whatever. The mind can no more contain bare representations than you can hang a shape on your wall without hanging a rectangle, circle or some other particular shape. It is sometimes but only rarely necessary to introduce an unspecified kind of representation. Doing so is always a mark of ignorance. We can see this by contrasting the comparatively rich predictions that can be made about infants’ abilities concerning actions, where conjectures about particular kinds

of representation are possible (recall Conjecture MP from Section 11.5), with the sparser predictions that can currently be made about mindreading because we do not yet know which kinds of representations underpin mindreading abilities.

Third, there is a three-fold distinction between formally, descriptively and explanatorily adequate (Section 3.5). When characterizing some infants’ abilities it is often useful to have a model of the domain. The model is often identified by a set of principles or a theory. When thinking about the model in relation to the infant’s abilities, we can regard it as merely formally adequate; that is, someone who used the model as a guide to reality, was otherwise omniscient and had unlimited cognitive resources could use it to manifest the ability. Or we may go further and state that the model is descriptively adequate; this is, it enables us to predict the extent and limits of the infants’ abilities. Nothing yet follows concerning representation. It is only in considering the explanatory adequacy of the model that we may need to invoke representation. Of course, things will quickly get messy if we try to make claims about representation on the basis of findings which speak only to formal or descriptive adequacy. But that is just a consequence of introducing theoretical claims which are not empirically motivated. None of this suggests there is anything especially problematic about the concept of representation.

17.8 Inferential and intentional isolation

Solving puzzles about the patterns of infants’ and adults’ successes and failures, and the associated linking problems, might enable us to describe what is going on in infants’ minds; but it leaves our overall question unanswered. The question was how humans first come to know simple facts about particular objects, actions and minds.

In attempting to answer this question, we are confronted by a dilemma concerning intentional isolation and lack of inferential integration.

On one horn, core knowledge (whatever exactly that turns out to be) must somehow contribute to the developmental acquisition of knowledge. It is tempting to think that this might involve some process of inference from core knowledge to knowledge, or at least an operation that transforms contents of core knowledge states into contents of knowledge states. Existing theories proceed along just these lines (Section 8.3). How else a transition from core knowledge to knowledge proper might occur remains a mystery.

On the other horn, core knowledge is invoked in part to solve puzzles about the patterns of infants’ and adults’ successes and failures (Section 17.3). This indicates (but does not demonstratively require, of course) that core knowledge is cut off from knowledge and other states: its influence on them must be limited in ways that preclude inferential integration. Otherwise the apparently puzzling patterns of performance could hardly persist into adulthood; and theories invoking core knowledge would generate the same incorrect predictions that versions of the Simple View do. Further, the effects of core knowledge may involve metacognitive feelings, behaviours and other intentional isolators. If, as I suggest, only intentional isolators can connect core knowledge to knowledge proper, then core knowledge is both inferentially and intentionally isolated from knowledge. This is inconsistent with the idea that a transition

from core knowledge to knowledge proper could involve an operation transforming the contents of one kind of state into (parts of) the contents of another.

My guess is that inferential and intentional isolation is real. Further, it is an architectural virtue and not a deficit. Any complex system with human-like capacities for error faces a trade-off between robustness and flexibility. Robustness requires a conservative attitude towards learning and change, whereas flexibility requires not being constrained by heuristics and assumptions in responding to new information, even at the risk of introducing large and potentially fatal errors. By having core systems isolated from knowledge systems, it is possible to make two different trade-offs simultaneously. If a philosopher can convince herself that tables, chairs and people are merely artefacts of her own mind, or even that the universe consists of herself alone, this will not directly impair her, thanks to inferential and intentional isolation.1009 The isolation allows core knowledge to be ultra conservative, ensuring a basic level of functioning in the world, while knowledge proper makes the complementary trade-off.

Understanding the developmental emergence of knowledge requires understanding how core knowledge contributes despite its isolation. Achieving takes us outside the infant’s head and into its social world.

17.9 Rediscovery is joint action

Dramatic discoveries about core knowledge are rarely thought about in connection with infants’ social skills and the role these play a role in cognitive development. This is understandable. After all, whatever the truth about innateness turns out to be, there is no indication in any of the discoveries we have considered that core knowledge is in any interesting sense a social phenomenon. This is a major challenge to any suggestion that all human cognitive development depends on social interaction in ways that development in other eusocial animals does not.

Reflection on inferential and intentional isolation motivates reconsideration. If the role of core knowledge in the emergence of knowledge proper can only be direct, what could core knowledge do? Perhaps core knowledge of objects, minds and actions enables infants not only to interact with physical objects but also to perform joint actions involving collective goals with the people around them. And perhaps these joint actions are somehow implicated in explaining how humans first acquire knowledge of simple facts about particular things.

This idea is so underspecified it is only just worth mentioning. To work it out in any detail, we would need to consider many further domains in which infants manifest abilities, including the moral, chromatic, numerical, spatial, communicative and linguistic. We would also to know much more about joint action in development. But we have made a step forwards in seeing the shape a story about the developmental emergence of knowledge might take. Core knowledge and joint action both need to feature in this story, and one role for core knowledge is to enable joint actions, which in turn allow humans to rediscover what is in some sense already encoded in their core knowledge.

Notes

Glossary

agent An agent of an action is a person or other individual who performs that action. When Ayesha kicks the door, she is the agent of the kick.

altercentric Interference The unintentional or non-purposive influence of another's belief (or other mental states) on your own belief (or corresponding mental states).

Assumption of Representational Connections The assumption that the transition from early-developing forms of representation to belief or knowledge involves operations on the early-developing forms of representation which transform their contents into (components of) the contents of knowledge states. The Assumption of Representational Connections is implicitly required by many theories about the developmental emergence of knowledge, but not by the view that development is rediscovery.

A-task Any false belief task that typically developing children tend to fail until around 3–5 years of age.

automatic In this book, a process is automatic just if whether or not it occurs is to a significant extent independent of your current task, motivations and intentions. To say that mindreading is automatic is to say that it involves only automatic processes. The term 'automatic' has been used in a variety of ways by other authors: see Moors (2014, 22) for a one-page overview, Moors and De Houwer (2006) for a detailed theoretical review, or Bargh (1992) for a classic and very readable introduction.

canonical model of minds and actions A model specified by a canonical theory of the mental.

canonical theory of the mental A theory of the mental which features attitudes such as belief, desire, knowledge and intention and relies on a system of propositions to distinguish their contents.

CLSTX Conjecture Four- and 5-month-olds' abilities to segment objects, to represent them while briefly unperceived and to track their causal interactions are not grounded on belief

or knowledge: instead they are consequences of the operations of a system of object indexes (Leslie et al. 1998; Scholl and Leslie 1999; Carey and Xu 2001; Scholl 2007).

collective goal A collective goal is any outcome to which the actions of two or more agents are collectively directed. And for their actions to be collectively directed to an outcome is for their actions to be directed to an outcome where this is not, or not only, a matter of each action being individually directed to that outcome.

Conjecture MP In the first nine months of life, all proper pure goal tracking is explained by the Motor Theory. Other pure goal-tracking processes emerge later in development. Further, a mere target-tracking process is also present in these infants. This process is identical to perceptual animacy in adults. And appearances that these infants' pure goal-tracking abilities are not limited by what they can represent motorically are misleading: they are due to mistaking mere target tracking for proper goal tracking.

Conjecture O Four- and 5-month-olds' abilities to segment objects, to represent them while briefly unperceived and to track their causal interactions are not grounded on belief or knowledge: instead they are consequences of the operations of a system of object indexes … and of a further, independent capacity to track physical objects which involves motor representations and processes.

Conjecture O^m This is Conjecture O together with the further conjecture that errors in operations on object indexes and motor representations can give rise to metacognitive feelings of surprise. Given the further hypothesis that these metacognitive feelings can explain looking behaviours on habituation and violation-of-expectation tasks, Conjecture O^m provides a solution to the Linking Problem which is consistent with the puzzling pattern of findings about 4-month-olds' abilities concerning physical objects summarized in Table 4.1.

contrast case A contrast case is a pair of events which are as similar as possible except that one is a joint action while the other is not.

core knowledge For an individual to have core knowledge concerning a domain such as physical objects, actions or minds is for her to have a core system specifically for this domain. For someone to have core knowledge of a particular principle or fact is for her to have a core system where either the core system includes a representation of that principle or else the principle plays a special role in describing the core system. Core knowledge is not knowledge, and you can have core knowledge of things that are untrue (for this reason, Carey (2009, 10) recommends the term 'core cognition' for states of core knowledge).

Core Knowledge View According to this view, the Principles of Object Perception are not things infants know but rather are encoded in their core knowledge. The operations of core system enable this core knowledge, together with perceptual inputs, to generate expectations concerning particular physical objects. And these expectations are not knowledge states but representations in core systems.

core system This book uses a non-standard, minimally informative notion of core system on which a 'core system' for a particular domain is simply whatever it is that underpins the earliest abilities infants manifest in that domain (see Section 8.2). This allows that core systems may lack uniformity across domains and unity within a domain: that is, different

kinds of system may qualify as 'core' in different domains, and a core system may comprise two or more largely distinct systems (see Section 6.9).

However, core systems are standardly identified by giving a list of features. The lists vary between researchers and times. Carey and Spelke (1996, 520) assert that core systems are largely innate, informationally encapsulated (that is, their operations are largely unaffected by things you know or believe, and by core knowledge in other core systems), largely unchanging over the course of development (so adults and infants alike have the same core systems). They also say that the inputs to core systems are the outputs of perceptual systems, so that architecturally core systems in human adults occupy a position between perception and knowledge. Finally, core systems are also held to arise from systems already present in the evolutionary ancestors of modern humans. Carey (2009) adds that the representations in core systems are iconic representations.

descriptively adequate Some principles are descriptively adequate to characterize a person's ability just if they enable us to predict the extent and limits of that person's ability. Contrasts with formally adequate and explanatorily adequate (see Table 3.2).

Developmental Motor Conjecture In the first nine months of life, all pure goal tracking is explained by the Motor Theory of Goal Tracking. Other goal-tracking processes emerge later in development.

dishabituation See habituation.

dual process theory Any theory concerning abilities in a particular domain on which those abilities involve two or more processes which are distinct in this sense: the conditions which influence whether one mindreading process occurs differ from the conditions which influence whether another occurs.

dual process theory of mindreading A theory on which mindreading involves two or more processes which are distinct in this sense: the conditions which influence whether one mindreading process occurs differ from the conditions which influence whether another occurs. (For background on dual process theories in social cognition generally, Sherman, Gawronski and Trope (2014) is an excellent collection of essays.)

explanatorily adequate For some principles to be explanatorily adequate to characterize a person's ability is for there to be a link between the principles and her mind in virtue which she has this ability. Where such a link exists, the principles do not merely describe her ability. Instead they play a role in specifying the nature of processes, representations or systems underpinning it as, for instance, when the principles specify the contents of knowledge states on which the ability is based. Contrasts with formally adequate and descriptively adequate (see Table 3.2).

formally adequate Some principles are formally adequate to characterize an ability to the extent that someone who took the principles to be true, was otherwise omniscient and had unlimited cognitive resources could use the principles to manifest the ability. Contrasts with descriptively adequate and explanatorily adequate (see Table 3.2).

goal A goal of an action is an outcome to which it is directed.

habituation Habituation is used to test hypotheses about which events are interestingly different to an infant. In a habituation experiment, infants are shown an event repeatedly until it no longer holds their interest, as measured by how long they look at it. The infants

are then divided into two (or more) groups and each group is shown a new event. How much longer do they look at the new event than at the most recent presentation of the old event? This difference in looking times indicates dishabituation, or the reawakening of interest. Given the assumption that greater dishabituation indicates that the old and new events are more interestingly different to the infant, evidence from patterns of dishabituation can sometimes support conclusions about patterns in how similar and different events are to infants.

iconic representation A representation is iconic if parts of the representation represent parts of the thing represented. Pictures are paradigm examples of iconic representations. For example, in a picture of a flower, some parts of the picture may represent petals while others represent the stem.

independent processes Two kinds of process are independent just if the conditions which influence whether a process of one kind will yield an incorrect response differ to some extent from the conditions that influence whether a process of the other kind will yield an incorrect response.

inferential integration For states to be inferentially integrated means that: (1) they can come to be non-accidentally related in ways that are approximately rational thanks to processes of inference and practical reasoning; and (2) in the absence of obstacles such as time pressure, distraction, motivations to be irrational, self-deception or exhaustion, approximately rational harmony will characteristically be maintained among those states that are currently active.

innate Not learned. While everyone disagrees about what innateness is (see Samuels 2004), in this book a cognitive ability is innate just if its developmental emergence is not a direct consequence of data-driven learning.

intentional isolation Two representations are intentionally isolated when the only links between them, if any, are provided by intentional isolators.

intentional isolator An event or state which links representations but either lacks intentional features entirely or else has intentional features that are only very distantly related to those of the two representations it links. Metacognitive feelings and behaviours are paradigm intentional isolators.

Joint Action Conjecture Abilities to perform joint actions play a role in explaining the developmental emergence of knowledge, including knowledge of others' minds.

knowledge proper Knowledge proper (or 'knowledge' for short) is constitutively linked to practical reasoning and to inference. Knowledge is the kind of thing that can typically influence how you act when you act purposively, and it is the kind of thing that can influence purposive actions in any domain at all. Knowledge is also the kind of thing that you can sometimes arrive at by inference and which can enable you to make new inferences in any domain at all. A state that is not linked to practical reasoning and to inference in these ways is not knowledge.

Linking Problem The problem of explaining what links the Principles of Object Perception to the abilities in virtue of which infants (and others) are able to segment objects, represent them as persisting and track their casual interactions (see Section 4.4). The Simple View, the CLSTX Conjecture and Conjecture O are competing attempts to solve it.

Marr's three levels Marr (1982, 22ff) distinguished three things a theory of a cognitive system should provide. The computational description specifies what the system is for and what it does in the broadest terms. Fixing representations and algorithms specifies how the system represents its inputs and outputs as well as how it transforms one into the other. Finally, to describe the hardware implementation is to specify how the representations and algorithms are physically realized.

metacognitive feeling Paradigm examples of metacognitive feelings include the feeling of familiarity, the feeling that something is on the tip of your tongue, the feeling of confidence and the feeling that someone's eyes are boring into your back. In this book, I propose that one characteristic of metacognitive feelings is that either they lack intentional objects altogether, or else what their subjects take them to be about is typically only very distantly related to their intentional objects. (This is controversial—see Dokic 2012, for a variety of conflicting theories.)

Mindreading Puzzle The puzzle is to determine which of these three collectively inconsistent claims are false. For many children, there is an age at which:

in performing false belief tasks which are A-tasks, the child relies on a model of minds and actions not incorporating beliefs;

in performing false belief tasks which are not A-tasks, such as tasks involving anticipatory looking or violation-of-expectation, the child relies on a model of minds and actions incorporating beliefs;

the child has a single model of minds and actions.

minimal model of minds and actions A model specified by a minimal theory of mind.

minimal theory of mind A theory of the mental in which: (1) mental states are assigned functional roles that can readily be codified; and, (2), the contents of mental states can be distinguished by things which, like locations, shapes and colours, can be held in mind using some kind of quality space or feature map.

model of minds and actions A model of minds and actions is a way mental aspects of the world could be. A model is not a theory, although it may be specified by one.

motor representation The kind of representation characteristically involved in preparing, performing and monitoring sequences of small-scale actions such as grasping, transporting and placing an object. They represent actual, possible, imagined or observed actions and their effects.

Motor Theory of Goal Tracking Tracking goals is acting in reverse (see Section 11.2 and Sinigaglia and Butterfill 2016).

non-A-task Any false belief task that typically developing children tend to pass in their first or second year of life.

object index An object index is a mental pointer to an object; it is the mental counterpart of the finger you might use to keep track of a moving object by pointing at it and following its path with your finger. See further Scholl (2001, 27ff).

object-specific preview benefit When a feature (for example, the letter 'T') presented earlier reappears now, how long will it take people to confirm that they saw this feature

earlier? All other things being equal, people are mostly faster when the feature reappears on the same object than when it reappears on a different object. This difference in response times is the object-specific preview benefit (Kahneman, Treisman and Gibbs 1992).

perceptual animacy The detection by broadly perceptual processes of animate objects (Scholl and Gao 2013, 201–2).

poverty of stimulus argument An argument used to establish that something is innate. A poverty of stimulus argument aims to establish that something humans acquire is not acquired by data driven-learning (see Pullum and Scholz 2002).

poverty of theory argument An inexpensive alternative to a poverty of stimulus argument.

Principles of Object Perception Principles which characterize how physical objects are perceived. These are thought to include no action at a distance, rigidity, boundedness and cohesion. See Table 3.1 for a brief statement of these principles.

pure goal tracking Tracking goals is pure when does not involve ascribing intentions or any other mental states.

rediscovery Rediscovery occurs when principles related to those already implicit in early-developing capacities to engage with physical objects, colours, mental states or some other domain must be discovered anew. Where there is rediscovery, there are no direct representational connections between the early-developing abilities and the beliefs or knowledge which emerge. Instead the early-developing abilities influence behaviour, guide attention and give rise to metacognitive feelings.

signature limit A signature limit of a system is a pattern of behaviour the system exhibits which is both defective, given what the system is for and peculiar to that system. A signature limit of a model is a set of predictions derivable from the model, which are incorrect, and which are not predictions of other models under consideration.

Simple View This term is used for two thematically related claims. Concerning physical objects, the Simple View is the claim that the Principles of Object Perception are things we know or believe, and we generate expectations from these principles by a process of inference. Concerning the goals of actions, the Simple View is the claim that the principles comprising the Teleological Stance are things we know or believe, and we are able to track goals by making inferences from these principles.

target The targets of an action (if any) are the things towards which it is directed.

Teleological Stance To adopt the Teleological Stance is to exploit certain principles concerning the optimality of goal-directed actions in tracking goals (for details, see Section 10.4).

track a belief For a process to track someone's belief that p is for it to non-accidentally depend in some way on whether she believes that p. For someone to track beliefs is for there to be processes in her which track some beliefs.

track a goal For a process to track a goal of an action is for how that process unfolds to non-accidentally depend in some way on whether that outcome is indeed a goal of the action. For someone to track the goals of an action is for there to be processes in her which track one or more goals of that action.

violation-of-expectation Violation-of-expectation experiments test hypotheses about what infants expect by comparing their responses to two events. The responses compared are usually looking durations. Looking durations are linked to infants' expectations by the assumption that, all things being equal, infants will typically look longer at something which violates an expectation of theirs than something which does not. Accordingly, with careful controls, it is sometimes possible to draw conclusions about infants' expectations from evidence that they generally look longer at one event than another.

Bibliography

Adolphs, Ralph. 2010. 'Conceptual Challenges and Directions for Social Neuroscience'. Neuron 65 (6): 752–67. http://dx.doi.org/10.1016/j.neuron.2010.03.006.

Adolphs, Ralph, Hanna Damasio, Daniel Tranel, Greg Cooper and Antonio R. Damasio. 2000. 'A Role for Somatosensory Cortices in the Visual Recognition of Emotion as Revealed by Three-Dimensional Lesion Mapping'. The Journal of Neuroscience 20 (7): 2683–90. Available at: www.jneurosci.org/content/20/7/2683.

Aguiar, Andréa and Renée Baillargeon. 2002. 'Developments in Young Infants' Reasoning About Occluded Objects'. Cognitive Psychology 45: 267–336.

Akhtar, Nameera, Maureen Callanan, Geoffrey K. Pullum and Barbara C. Scholz. 2004. 'Learning Antecedents for Anaphoric One'. Cognition 93 (2): 141–45. doi:10.1016/j.cognition.2003.12.002.

Alonso, Facundo M. 2009. 'Shared Intention, Reliance and Interpersonal Obligations'. Ethics 119 (3): 444–75. doi:10.1086/599984.

Alvarez, George A. and Steven L. Franconeri. 2007. 'How Many Objects Can You Track?: Evidence for a Resource-Limited Attentive Tracking Mechanism'. Journal of Vision 7 (13): 14. doi:10.1167/7.13.14.

Ambrosini, Ettore, Marcello Costantini and Corrado Sinigaglia. 2011. 'Grasping with the Eyes'. Journal of Neurophysiology 106 (3): 1437–42. doi:10.1152/jn.00118.2011.

Ambrosini, Ettore, Vasudevi Reddy, Annette de Looper, Marcello Costantini, Beatriz Lopez and Corrado Sinigaglia. 2013. 'Looking Ahead: Anticipatory Gaze and Motor Ability in Infancy'. PLOS One 8 (7): e67916. doi:10.1371/journal.pone.0067916.

Ambrosini, Ettore, Corrado Sinigaglia and Marcello Costantini. 2012. 'Tie My Hands, Tie My Eyes'. Journal of Experimental Psychology: Human Perception and Performance 38 (2): 263–6. doi:10.1037/a0026570.

Apperly, Ian A., Elisa Back, Dana Samson and Lisa France. 2008. 'The Cost of Thinking About False Beliefs: Evidence from Adults' Performance on a Non-Inferential Theory of Mind Task'. Cognition 106: 1093–1108.

Apperly, Ian A., Kevin Riggs, A. Simpson, C. Chiavarino and Dana Samson. 2006. 'Is Belief Reasoning Automatic?' Psychological Science 17 (10): 841–44.

Apperly, Ian A., Dana Samson and Glyn W. Humphreys. 2009. 'Studies of Adults Can Inform Accounts of Theory of Mind Development'. Developmental Psychology 45 (1): 190–201.

Apperly, Ian A., Frances Warren, Benjamin J. Andrews, Jay Grant and Sophie Todd. 2011. 'Developmental Continuity in Theory of Mind: Speed and Accuracy of Belief-Desire Reasoning in Children and Adults'. Child Development 82 (5): 1691–703. doi:10.1111/j.1467–8624.2011.01635.x.

Aslin, Richard N. 2007. 'What's in a Look?' Developmental Science 10 (1): 48–53. doi:10.1111/j.1467–7687.2007.00563.x.

Astington, Janet and Alison Gopnik. 1988. 'Knowing You've Changed Your Mind: Children's Understanding of Representational Change'. In Developing Theories of Mind, edited by Janet Astington, Paul Harris and David Olson. Cambridge: Cambridge University Press.

Astington, Paul Harris, and David Olson. 1991. “Developing Understanding of Desire and Intention.” In Natural Theories of the Mind: Evolution, Development and Simulation of Everyday Mindreading, edited by Andrew Whiten, 39–50. Oxford: Blackwell.

Babinsky, Erin, Oliver Braddick and Janette Atkinson. 2011. 'Infants and Adults Reaching in the Dark'. Experimental Brain Research 217 (2): 237–49. doi:10.1007/s00221-011-2984–5.

Back, E. and Ian A. Apperly. 2010. 'Two Sources of Evidence on the Non-Automaticity of True and False Belief Ascription'. Cognition 115 (1): 54–70.

Baillargeon, Renée. 1987. 'Object Permanence in 3.5- and 4.5-Month-Old Infants'. Developmental Psychology 23 (5): 655–64.

Baillargeon, Renée. 2001. 'Infants' Physical Knowledge: Of Acquired Expectations and Core Principles'. In Language, Brain, and Cognitive Development: Essays in Honor of Jacques Mehler, edited by Emmanuel Dupoux, 341–61. Cambridge, MA: MIT Press.

Baillargeon, Renée. 2002. 'The Acquisition of Physical Knowledge in Infancy: A Summary in Eight Lessons'. In Blackwell Handbook of Childhood Cognitive Development, edited by Usha Goswami, 47–83. Oxford: Blackwell.

Baillargeon, Renée, Rose M. Scott and Zijing He. 2010. 'False-Belief Understanding in Infants'. Trends in Cognitive Sciences 14 (3): 110–18.

Baillargeon, Renée, Rose M. Scott, Zijing He, Stephanie Sloane, Peipei Setoh, Kyong-sun Jin, Di Wu and Lin Bian. 2015. 'Psychological and Sociomoral Reasoning in Infancy'. In APA Handbook of Personality and Social Psychology, Vol. 1: Attitudes and Social Cognition, edited by M. Mikulincer, P. R. Shaver, E. Borgida and J. A. Bargh, 79–150. Washington, DC: American Psychological Association.

Bartsch, Karen and Henry M. Wellman. 1995. Children Talk About the Mind. Oxford: Oxford University Press.

Behne, Tanya, Malinda Carpenter, Josep Call and Michael Tomasello. 2005. 'Unwilling Versus Unable: Infants' Understanding of Intentional Action'. Developmental Psychology 41 (2): 328–37.

Behne, Tanya, Malinda Carpenter and Michael Tomasello. 2005. 'One-Year-Olds Comprehend the Communicative Intentions Behind Gestures in a Hiding Game'. Developmental Science 8 (6): 492–9.

Bennett, Jonathan. 1976. Linguistic Behaviour. Cambridge: Cambridge University Press.

Benson, Jeannette E., Mark A. Sabbagh, Stephanie M. Carlson and Philip David Zelazo. 2013. 'Individual Differences in Executive Functioning Predict Preschoolers' Improvement from Theory-of-Mind Training'. Developmental Psychology 49 (9): 1615–27. doi:10.1037/a0031056.

Bermúdez, José Luis. 2003. Thinking Without Words. Oxford: Oxford University Press.

Bertenthal, Bennett I., Gustaf Gredebäck and Ty W. Boyer. 2013. 'Differential Contributions of Development and Learning to Infants' Knowledge of Object Continuity and Discontinuity'. Child Development 84 (2): 413–21. doi:10.1111/cdev.12005.

Berthier, N. E., S. De Blois, C. R. Poirier, M. A. Novak and R. K Clifton. 2000. 'Where's the Ball? Two- and Three-Year-Olds Reason About Unseen Events'. Developmental Psychology 36 (3): 394–401.

Berwick, Robert C., Paul Pietroski, Beracah Yankama and Noam Chomsky. 2011. 'Poverty of the Stimulus Revisited'. Cognitive Science 35 (7): 1207–42. doi:10.1111/j.1551–6709.2011.01189.x.

Blomberg, Olle. 2016. 'Common Knowledge and Reductionism About Shared Agency'. Australasian Journal of Philosophy 94 (2): 315–26. doi:10.1080/00048402.2015.1055581.

Boghossian, Paul A. 2003. 'The Normativity of Content'. Philosophical Issues 13 (1): 31–45. doi:10.1111/1533–6077.00003.

Braddon-Mitchell, David and Frank Jackson. 1996. Philosophy of Mind and Cognition. Oxford: Blackwell.

Bratman, Michael E. 1987. Intentions, Plans, and Practical Reasoning. Cambridge, MA: Harvard University Press.

Bratman, Michael E.1992. 'Shared Cooperative Activity'. The Philosophical Review 101 (2): 327–41.

Bratman, Michael E. 1993. 'Shared Intention'. Ethics 104: 97–113.

Bratman, Michael E. 1997. 'I Intend That We J'. In Contemporary Action Theory, Vol. 2: Social Action, edited by Raimo Tuomela and Ghita Holmstrom-Hintikka. Dordrecht: Kluwer.

Bratman, Michael E. 2006. 'Dynamics of Sociality'. Midwest Studies in Philosophy 30: 1–15.

Bratman, Michael E. 2009. 'Modest Sociality and the Distinctiveness of Intention'. Philosophical Studies 144 (1): 149–65.

Bratman, Michael E. 2010. 'Agency, Time, and Sociality'. Proceedings and Addresses of the American Philosophical Association 84 (2): 7–26.

Bratman, Michael E. 2014. Shared Agency: A Planning Theory of Acting Together. Oxford: Oxford University Press.

Bratman, Michael E. 2015. 'Shared Agency: Replies to Ludwig, Pacherie, Petersson, Roth, and Smith'. Journal of Social Ontology 1 (1): 59–76.

Bremner, J. Gavin, Alan M. Slater and Scott P. Johnson. 2015. 'Perception of Object Persistence: The Origins of Object Permanence in Infancy'. Child Development Perspectives 9 (1): 713. doi:10.1111/cdep.12098.

Brown, Alan S. 2003. 'A Review of the Déjà Vu Experience'. Psychological Bulletin 129 (3): 394–413. doi:10.1037/0033–2909.129.3.394.

Brownell, Celia A. 2011. 'Early Developments in Joint Action'. Review of Philosophy and Psychology 2: 193–211. doi:10.1007/s13164–011–0056–1.

Brownell, Celia A., Geetha B. Ramani and Stephanie Zerwas. 2006. 'Becoming a Social Partner with Peers: Cooperation and Social Understanding in One- and Two-Year-Olds'. Child Development 77 (4): 803–21.

Bruderer, Alison G., D. Kyle Danielson, Padmapriya Kandhadai and Janet F. Werker. 2015. 'Sensorimotor Influences on Speech Perception in Infancy'. Proceedings of the National Academy of Sciences 112 (44): 13531–6. doi:10.1073/pnas.1508631112.

Buetti, Simona, Alejandro Lleras and Cathleen M. Moore. 2014. 'The Flanker Effect Does Not Reflect the Processing of “Task-Irrelevant” Stimuli: Evidence from Inattentional Blindness'. Psychonomic Bulletin & Review 21 (5): 1231–7. doi:10.3758/s13423–014–0602–9.

Bull, R., L. H. Phillips and C. A. Conway. 2008. 'The Role of Control Functions in Mentalizing: Dual-Task Studies of Theory of Mind and Executive Function'. Cognition 107 (2): 663–72.

Butler, Samantha C., N. E. Berthier and R. K Clifton. 2002. 'Two-Year-Olds' Search Strategies and Visual Tracking in a Hidden Displacement Task'. Developmental Psychology 38 (4): 581–90.

Buttelmann, David, Renée Baillargeon and Victoria Southgate. n.d. 'Interpreting Failed Replications of Early False-Belief Findings: Methodological and Theoretical Considerations'. Cognitive Development. doi:10.1016/j.cogdev.2018.06.001.

Buttelmann, Frances and Ágnes Melinda Kovács. n.d. '14-Month-Olds Anticipate Others' Actions Based on Their Belief About an Object's Identity'. Infancy. doi:10.1111/infa.12303.

Buttelmann, Frances, Janina Suhrke and David Buttelmann. 2015. 'What You Get Is What You Believe: Eighteen-Month-Olds Demonstrate Belief Understanding in an Unexpected-Identity Task'. Journal of Experimental Child Psychology 131 (March): 94–103. doi:10.1016/j.jecp.2014.11.009.

Butterfill, Stephen A. 2001. 'Awareness of Belief'. In Argument & Analyse: Ausgewählte Sektionsvorträge des 4. Internationalen Kongresses der Gesellschaft für analytische Philosophie, edited by Ansgar Beckermann and Christian Nimtz. Vol. 2. Bielefeld: Mentis.

Butterfill, Stephen A. 2009. 'Seeing Causes and Hearing Gestures'. Philosophical Quarterly 59 (236): 405–28.

Butterfill, Stephen A. 2012a. 'Interacting Mindreaders'. Philosophical Studies 165 (3): 841–63. doi:10.1007/s11098–012–9980-x.

Butterfill, Stephen A. 2012b. 'Joint Action and Development'. Philosophical Quarterly 62 (246): 23–47.

Butterfill, Stephen A. 2015. 'Perceiving Expressions of Emotion: What Evidence Could Bear on Questions About Perceptual Experience of Mental States?' Consciousness and Cognition 36: 438–51. doi:10.1016/j.concog.2015.03.008.

Butterfill, Stephen A. and Ian A. Apperly. 2013. 'How to Construct a Minimal Theory of Mind'. Mind and Language 28 (5): 606–37.

Butterfill, Stephen A. and Ian A. Apperly. 2016. 'Is Goal Ascription Possible in Minimal Mindreading?' Psychological Review 123 (2): 228–33. doi:10.1037/rev0000022.

Byrne, Alex. 2001. 'Intentionalism Defended'. The Philosophical Review 110 (2): 199–240.

Call, Josep and Michael Tomasello. 1999. 'A Nonverbal False Belief Task: The Performance of Children and Great Apes'. Child Development 70 (2): 381–95.

Campbell, J. (2002). Reference and Consciousness. Oxford: Oxford University Press.

Cannon, Erin N. and Amanda L. Woodward. 2012. 'Infants Generate Goal-Based Action Predictions'. Developmental Science 15 (2): 292–98. doi:10.1111/j.1467–7687.2011.01127.x.

Cardellicchio, Pasquale, Corrado Sinigaglia and Marcello Costantini. 2011. 'The Space of Affordances: A TMS Study'. Neuropsychologia 49 (5): 1369–72. doi:10.1016/j.neuropsychologia.2011.01.021.

Carey, Susan. 1995. 'On the Origin of Causal Understanding'. In Causal Cognition: A Multidisciplinary Debate, edited by Dan Sperber, David Premack and Ann James Premack. Oxford: Clarendon.

Carey, Susan. 2009. The Origin of Concepts. Oxford: Oxford University Press.

Carey, Susan and Elizabeth Spelke. 1994. 'Domain-Specific Knowledge and Conceptual Change'. In Mapping the Mind: Domain Specificity in Cognition and Culture, edited by Lawrence Hirschfeld and Susan Gelman. Cambridge: Cambridge University Press.

Carey, Susan and Elizabeth Spelke. 1996. 'Science and Core Knowledge'. Philosophy of Science 63: 515–33.

Carey, Susan and Fei Xu. 2001. 'Infants' Knowledge of Objects: Beyond Object Files and Object Tracking'. Cognition 80: 179–213.

Carlson, Stephanie M, Louis J Moses and Laura J Claxton. 2004. 'Individual Differences in Executive Functioning and Theory of Mind: An Investigation of Inhibitory Control and Planning Ability'. Journal of Experimental Child Psychology 87 (4): 299–319. doi:10.1016/j.jecp.2004.01.002.

Carpenter, Malinda. 2009. 'Just How Joint Is Joint Action in Infancy?' Topics in Cognitive Science 1 (2): 380–92. doi:10.1111/j.1756–8765.2009.01026.x.

Carpenter, Malinda, Josep Call and Michael Tomasello. 2002. 'A New False Belief Test for 36-Month-Olds'. British Journal of Developmental Psychology 20: 393–420.

Carruthers, Peter. 2013. 'Mindreading in Infancy'. Mind & Language 28 (2): 141–72. doi:10.1111/mila.12014.

Carruthers, Peter. 2015a. 'Two Systems for Mindreading?' Review of Philosophy and Psychology 7 (1): 141–62. doi:10.1007/s13164–015–0259-y.

Carruthers, Peter. 2015b. 'Mindreading in Adults: Evaluating Two-Systems Views'. Synthese (June): 1–16. doi:10.1007/s11229–015–0792–3.

Cassidy, Kimberly Wright, Deborah Shaw Fineberg, Kimberly Brown and Alexis Perkins. 2005. 'Theory of Mind May Be Contagious, but You Don't Catch It from Your Twin'. Child Development 76 (1): 97–106. doi:10.1111/j.1467–8624.2005.00832.x.

Chandler, Michael, Anna Fritz and Suzanne Hala. 1989. 'Small Scale Deceit: Deception as a Marker of Two-, Three-, and Four-Year-Olds' Early Theories of Mind'. Child Development 60: 1263–77.

Charles, Eric P. and Susan M. Rivera. 2009. 'Object Permanence and Method of Disappearance: Looking Measures Further Contradict Reaching Measures'. Developmental Science 12 (6): 991–1006. doi:10.1111/j.1467–7687.2009.00844.x.

Cheries, Erik W., Karen Wynn and Brian J. Scholl. 2006. 'Interrupting Infants' Persisting Object Representations: An Object-Based Limit?' Developmental Science 9 (5): F50–F58. doi:10.1111/j.1467–7687.2006.00521.x.

Chiandetti, Cinzia and Giorgio Vallortigara. 2011. 'Intuitive Physical Reasoning About Occluded Objects by Inexperienced Chicks'. Proceedings of the Royal Society B: Biological Sciences 278 (1718): 2621–7. doi:10.1098/rspb.2010.2381.

Choi, Hoon and Brian J. Scholl. 2006. 'Measuring Causal Perception: Connections to Representational Momentum?' Acta Psychologica 123 (12): 91–111. doi:10.1016/j.actpsy.2006.06.001.

Chomsky Noam. 1957. Syntactic Structures. Mansfield Centre, CT: Martino Fine Books.

Chomsky Noam. 1965. Aspects of the Theory of Syntax. Cambridge, MA: MIT Press.

Chomsky, Noam. 1995. 'Language and Nature'. Mind 104 (413): 1–61.

Christensen, Wayne and John Michael. 2016. 'From Two Systems to a Multi-Systems Architecture for Mindreading'. New Ideas in Psychology 40: 48–64. doi:10.1016/j.newideapsych.2015.01.003.

Clements, Wendy and Josef Perner. 1994. 'Implicit Understanding of Belief'. Cognitive Development 9: 377–95.

Cohen, Adam S. and Tamsin C. German. 2009. 'Encoding of Others' Beliefs Without Overt Instruction'. Cognition 111 (3): 356–63. doi:10.1016/j.cognition.2009.03.004.

Cohen, L. Jonathan. 1992. An Essay on Belief and Acceptance. Oxford: Clarendon Press.

Costantini, Marcello, Ettore Ambrosini, Pasquale Cardellicchio and Corrado Sinigaglia. 2014. 'How Your Hand Drives My Eyes'. Social Cognitive and Affective Neuroscience 9 (5): 705–11. doi:10.1093/scan/nst037.

Costantini, Marcello, Ettore Ambrosini, Corrado Sinigaglia and Vittorio Gallese. 2011. 'Tool-Use Observation Makes Far Objects Ready-to-Hand'. Neuropsychologia 49 (9): 2658–63. doi:10.1016/j.neuropsychologia.2011.05.013.

Costantini, Marcello, Ettore Ambrosini, Gaetano Tieri, Corrado Sinigaglia and Giorgia Committeri. 2010. 'Where Does an Object Trigger an Action? An Investigation About Affordances in Space'. Experimental Brain Research 207 (1–2): 95–103. doi:10.1007/s00221–010–2435–8.

Crivello, Cristina and Diane Poulin-Dubois. 2017. 'Infants' False Belief Understanding: A Non-Replication of the Helping Task'. Cognitive Development, November. doi:10.1016/j.cogdev.2017.10.003.

Csibra, Gergely. 2003. 'Teleological and Referential Understanding of Action in Infancy'. Philosophical Transactions: Biological Sciences 358 (1431): 447–58.

Csibra, Gergely. 2008. 'Action Mirroring and Action Understanding: An Alternative Account'. Sensorimotor Foundations of Higher Cognition. Attention and Performance XXII: 435–59.

Csibra, Gergely and György Gergely. 1998. 'The Teleological Origins of Mentalistic Action Explanations: A Developmental Hypothesis'. Developmental Science 1 (2): 255–9.

Csibra, Gergely and György Gergely. 2009. 'Natural Pedagogy'. Trends in Cognitive Sciences 13 (4): 148–53.

Csibra, Gergely and György Gergely. 2013. 'Teleological Understanding of Actions'. In Navigating the Social World: What Infants, Children, and Other Species Can Teach Us, edited by M. R. Banaji and Susan A. Gelman, 37–43. Oxford: Oxford University Press.

Custer, Wendy. 1996. 'A Comparison of Young Children's Understanding of Contradictory Representations in Pretense, Memory, and Belief'. Child Development 67: 678–88.

Daum, Moritz M., Manja Attig, Ronald Gunawan, Wolfgang Prinz and Gustaf Gredeback. 2012. 'Actions Seen Through Babies' Eyes: A Dissociation Between Looking Time and Predictive Gaze'. Frontiers in Psychology 3 (September). doi:10.3389/fpsyg.2012.00370.

Davidson, Donald. 1963. 'Actions, Reasons and Causes'. In Essays on Actions and Events. Oxford: Oxford University Press.

Davidson, Donald. 1967. 'The Logical Form of Action Sentences'. In Essays on Actions and Events. Oxford: Oxford University Press.

Davidson, Donald. 1971. 'Agency'. In Agent, Action, and Reason, edited by Robert Binkley, Richard Bronaugh and Ausonia Marras, 3–25. Toronto: University of Toronto Press.

Davidson, Donald. 1976. 'Hume's Cognitive Theory of Pride'. The Journal of Philosophy 73 (19): 744–57.

Davidson, Donald. 1995. 'The Problem of Objectivity'. Tijdshrift Voor Filosofie 57 (2): 203–20.

Davidson, Donald. 1997. 'Seeing Through Language'. In Thought and Language, edited by John Preston. Royal Institute of Philosophy Supplement, Vol. 42. Oxford: Oxford University Press.

Davidson, Donald. 1999a. 'Replies to Critics'. In The Philosophy of Donald Davidson, edited by Lewis Edwin Hahn. Chicago: Open Court.

Davidson, Donald. 1999b. 'The Emergence of Thought'. Erkenntnis 51: 7–17.

Davidson, Donald. 2001. Subjective, Intersubjective, Objective. Oxford: Clarendon Press.

Davies, Martin. 1989. 'Tacit Knowledge and Subdoxastic States'. In Reflections on Chomsky, edited by Alexander George, 131–52. Oxford: Blackwell.

de Haan, Michelle. 2002. 'The Neuropsychology of Face Processing During Infancy and Childhood'. In Handbook of Developmental Cognitive Neuroscience, edited by Charles A. Nelson and Monica Luciana, 381–98. Cambridge, MA: MIT Press.

Deppe, Anja M., Patricia C. Wright and William A. Szelistowski. 2009. 'Object Permanence in Lemurs'. Animal Cognition 12 (2): 381–8. doi:10.1007/s10071–008–0197–5.

de Villiers, Peter A. de and Jill G. de Villiers. 2012. 'Deception Dissociates from False Belief Reasoning in Deaf Children: Implications for the Implicit Versus Explicit Theory of Mind Distinction'. British Journal of Developmental Psychology 30 (1): 188–209. doi:10.1111/j.2044–835X.2011.02072.x.

Devine, Rory T. and Claire Hughes. 2014. 'Relations Between False Belief Understanding and Executive Function in Early Childhood: A Meta-Analysis'. Child Development 85 (5): 1777–94. doi:10.1111/cdev.12237.

Dienes, Zoltán and Josef Perner. 1999. 'A Theory of Implicit and Explicit Knowledge'. Behavioral and Brain Sciences 22: 735–808.

Dittrich, Winand H. and Stephen Lea. 1994. 'Visual Perception of Intentional Motion'. Perception 23 (3): 253–68. doi:10.1068/p230253.

Doherty, Martin J. 1999. 'Selecting the Wrong Processor: A Critique of Leslie's Theory of Mind Mechanism Selection Processor Theory'. Developmental Science 2 (1): 81–85. doi:10.1111/1467–7687.00058.

Doherty, Martin J. 2011. 'A Two Systems Theory of Social Cognition: Engagement and Theory of Mind'. In Perception, Causation, and Objectivity, edited by Johannes Roessler, Hemdat Lerman and Naomi Eilan, 305–23. Oxford: Oxford University Press.

Dokic, Jérôme. 2012. 'Seeds of Self-Knowledge: Noetic Feelings and Metacognition'. In Foundations of Metacognition, edited by Michael J. Beran, Johannes L. Brandl, Josef Perner and Joëlle Proust, 302–21. Oxford University Press Oxford, England.

Dokic, Jérôme. 2016. 'Aesthetic Experience as a Metacognitive Feeling? A Dual-Aspect View'. Proceedings of the Aristotelian Society 116 (1): 69–88. doi:10.1093/arisoc/aow002.

Dörrenberg, Sebastian, Hannes Rakoczy and Ulf Liszkowski. 2018. 'How (Not) to Measure Infant Theory of Mind: Testing the Replicability and Validity of Four Non-Verbal Measures'. Cognitive Development, February. doi:10.1016/j.cogdev.2018.01.001.

Dretske, Fred. 2000. Perception, Knowledge and Belief. Cambridge: Cambridge University Press.

Durand, Karine and Roger Lécuyer. 2002. 'Object Permanence Observed in 4-Month-Old Infants with a 2D Display'. Infant Behavior and Development 25 (3): 269–78. doi:10.1016/S0163–6383(02)00098-X.

Edwards, Katheryn and Jason Low. 2017. 'Reaction Time Profiles of Adults' Action Prediction Reveal Two Mindreading Systems'. Cognition 160: 1–16. doi:10.1016/j.cognition.2016.12.004.

Edwards, Katheryn and Jason Low. 2019. 'Level 2 Perspective-Taking Distinguishes Automatic and Non-Automatic Belief-Tracking'. Cognition 193 (December): 104017. doi:10.1016/j.cognition.2019.104017.

El Kaddouri, Rachida, Lara Bardi, Diana De Bremaeker, Marcel Brass and Jan R. Wiersema. 2019. 'Measuring Spontaneous Mentalizing with a Ball Detection Task: Putting the Attention-Check Hypothesis by Phillips and Colleagues (2015) to the Test'. Psychological Research, April. doi:10.1007/s00426–019–01181–7.

Endedijk, Hinke M., Veronica C. O. Ramenzoni, Ralf F. A. Cox, Antonius H. N. Cillessen, Harold Bekkering and Sabine Hunnius. 2015. 'Development of Interpersonal Coordination Between Peers During a Drumming Task'. Developmental Psychology 51 (5): 714–21. doi:10.1037/a0038980.

Eriksen, Barbara A. and Charles W. Eriksen. 1974. 'Effects of Noise Letters Upon the Identification of a Target Letter in a Nonsearch Task'. Perception & Psychophysics 16 (1): 143–49. doi:10.3758/BF03203267.

Eshuis, Rik, Kenny R. Coventry and Mila Vulchanova. 2009. 'Predictive Eye Movements Are Driven by Goals, Not by the Mirror Neuron System'. Psychological Science 20 (4): 438–40.

Evans, Jonathan St. B. T. and Keith Frankish. 2009. In Two Minds: Dual Processes and Beyond. Oxford: Oxford University Press Oxford.

Falck-Ytter, Terje, Gustaf Gredeback and Claes von Hofsten. 2006. 'Infants Predict Other People's Action Goals'. Nature Neuroscience 9 (7): 878–9.

Fawcett, Christine and Ulf Liszkowski. 2012. 'Observation and Initiation of Joint Action in Infants'. Child Development 83 (2): 434–41. doi:10.1111/j.1467–8624.2011.01717.x.

Fiset, Sylvain and Vickie Plourde. 2013. 'Object Permanence in Domestic Dogs (Canis Lupus Familiaris) and Gray Wolves (Canis Lupus)'. Journal of Comparative Psychology 127 (2): 115–27. doi:10.1037/a0030595.

Fizke, Ella, Stephen A. Butterfill, Lea van de Loo, Eva Reindl and Hannes Rakoczy. 2017. 'Are There Signature Limits in Early Theory of Mind?' Journal of Experimental Child Psychology 162: 209–224.

Flanagan, J. Randall and Roland S. Johansson. 2003. 'Action Plans Used in Action Observation'. Nature 424 (6950): 769–71.

Flavell, John H. 1963. The Developmental Psychology of Jean Piaget. Princeton, NJ: D. Van Nostrand.

Flombaum, Jonathan I. and Brian J. Scholl. 2006. 'A Temporal Same-Object Advantage in the Tunnel Effect: Facilitated Change Detection for Persisting Objects'. Journal of Experimental Psychology. Human Perception and Performance 32 (4): 840–53. doi:10.1037/0096–1523.32.4.840.

Fodor, Jerry. 1983. The Modularity of Mind: An Essay on Faculty Psychology.. Cambridge, MA: MIT Press.

Foster, Meadhbh I. and Mark T. Keane. 2015. 'Why Some Surprises Are More Surprising Than Others: Surprise as a Metacognitive Sense of Explanatory Difficulty'. Cognitive Psychology 81 (September): 74–116. doi:10.1016/j.cogpsych.2015.08.004.

Franconeri, Steven L., Zenon W. Pylyshyn and Brian J. Scholl. 2012. 'A Simple Proximity Heuristic Allows Tracking of Multiple Objects Through Occlusion'. Attention, Perception, & Psychophysics 74 (4): 691–702. doi:10.3758/s13414–011–0265–9.

Furlanetto, T., Cristina Becchio, Dana Samson and Ian A. Apperly. 2015. 'Altercentric Interference in Level-1 Visual Perspective-Taking Reflects the Ascription of Mental States, Not Sub-Mentalizing'. Journal of Experimental Psychology: Human Perception and Performance 42 (2).

Gao, Tao, George E. Newman and Brian J. Scholl. 2009. 'The Psychophysics of Chasing: A Case Study in the Perception of Animacy'. Cognitive Psychology 59 (2): 154–79. doi:10.1016/j.cogpsych.2009.03.001.

Garnham, Wendy and Josef Perner. 2001. 'Actions Really Do Speak Louder Than Words—but Only Implicitly: Young Children's Understanding of False Belief in Action'. British Journal of Developmental Psychology 19: 413–32.

Garnham, Wendy and Ted Ruffman. 2001. 'Doesn't See, Doesn't Know: Is Anticipatory Looking Really Related to Understanding or Belief?' Developmental Science 4 (1): 94–100.

Gergely, György and Gergely Csibra. 2003. 'Teleological Reasoning in Infancy: The Naive Theory of Rational Action'. Trends in Cognitive Sciences 7 (7): 287–92.

Gergely, György, Z. Nadasky, Gergely Csibra and S. Biro. 1995. 'Taking the Intentional Stance at 12 Months of Age'. Cognition 56: 165–93.

Gerson, Sarah A., Harold Bekkering and Sabine Hunnius. 2016. 'Social Context Influences Planning Ahead in Three-Year-Olds'. Cognitive Development 40: 120–31. doi:10.1016/j.cogdev.2016.08.010.

Gilbert, Margaret P. 1990. 'Walking Together: A Paradigmatic Social Phenomenon'. Midwest Studies in Philosophy 15: 1–14.

Gilbert, Margaret P. 2006. 'Rationality in Collective Action'. Philosophy of the Social Sciences 36 (1): 3–17.

Gilbert, Margaret P. 2013. Joint Commitment: How We Make the Social World. Oxford: Oxford University Press.

Godfrey-Smith, Peter. 2005. 'Folk Psychology as a Model'. Philosophers' Imprint 5 (6).

Gold, Natalie and Robert Sugden. 2007. 'Collective Intentions and Team Agency'. Journal of Philosophy 104 (3): 109–37.

Gómez, Juan-Carlos. 2005. 'Species Comparative Studies and Cognitive Development'. Trends in Cognitive Sciences 9 (3): 118–25. doi:10.1016/j.tics.2005.01.004.

Gopnik, Alison. 1996. 'The Scientist as Child'. Philosophy of Science 63: 485–514.

Gopnik, Alison and Janet Astington. 1988. 'Children's Understanding of Representational Change and Its Relation to the Understanding of False Belief and the Appearance-Reality Distinction'. Child Development 59: 26–37.

Gopnik, Alison and Andrew N. Meltzoff. 1997. Words, Thoughts, and Theories: Learning, Development, and Conceptual Change. Cambridge, MA: MIT Press.

Gopnik, Alison and Virginia Slaughter. 1991. 'Young Children's Understanding of Changes in Their Mental States'. Child Development 62: 98–110.

Gopnik, Alison, Virginia Slaughter and Andrew N. Meltzoff. 1994. ‘Changing Your Views: How Understanding Visual Perception Can Lead to a New Theory of Mind’. In Children’s Early Understanding of Mind: Origins and Development, edited by Charlie Lewis and Peter Mitchell. Hove: Erlbaum.

Gordon, Robert. 2000. ‘Sellars’s Ryleans Revisited’. Protosociology 14: 102–14.

Gräfenhain, Maria, Tanya Behne, Malinda Carpenter and Michael Tomasello. 2009. ‘Young Children’s Understanding of Joint Commitments’. Developmental Psychology 45 (5): 1430–43.

Gredebäck, Gustaf and Moritz M. Daum. 2015. ‘The Microstructure of Action Perception in Infancy: Decomposing the Temporal Structure of Social Information Processing’. Child Development Perspectives 9 (2): 79–83. doi:10.1111/cdep.12109.

Gredebäck, Gustaf and Terje Falck-Ytter. 2015. ‘Eye Movements During Action Observation’. Perspectives on Psychological Science 10 (5): 591–8. doi:10.1177/1745691615589103.

Gredebäck, Gustaf and Annika Melinder. 2010. ‘Infants’ Understanding of Everyday Social Interactions: A Dual Process Account’. Cognition 114 (2): 197–206. doi:10.1016/j.cognition.2009.09.004.

Gredebäck, Gustaf and Annika Melinder. 2011. ‘Teleological Reasoning in 4-Month-Old Infants: Pupil Dilations and Contextual Constraints’. PLOS ONE 6 (10): e26487. doi:10.1371/journal.pone.0026487.

Green, Dorota, Qi Li, Jeffrey J. Lockman and Gustaf Gredebäck. 2016. ‘Culture Influences Action Understanding in Infancy: Prediction of Actions Performed with Chopsticks and Spoons in Chinese and Swedish Infants’. Child Development 87 (3): 736–46. doi:10.1111/cdev.12500.

Grosse Wiesmann, Charlotte, Angela D. Friederici, Tania Singer and Nikolaus Steinbeis. 2016. ‘Implicit and Explicit False Belief Development in Preschool Children’. Developmental Science 20 (5).

Haith, Marshall. 1998. ‘Who Put the Cog in Infant Cognition? Is Rich Interpretation Too Costly?’ Infant Behavior and Development 21 (2): 167–79.

He, Zijing, Matthias Bolz and Renée Baillargeon. 2011. ‘False-Belief Understanding in 2.5-Year-Olds: Evidence from Violation-of-Expectation Change-of-Location and Unexpected-Contents Tasks’. Developmental Science 14 (2): 292–305.

He, Zijing, Matthias Bolz and Renée Baillargeon. 2012. ‘2.5-Year-Olds Succeed at a Verbal Anticipatory-Looking False-Belief Task’. British Journal of Developmental Psychology 30 (1): 14–29. doi:10.1111/j.2044-835X.2011.02070.x.

Heal, Jane. 2002. ‘On First-Person Authority’. Proceedings of the Aristotelian Society 102: 1–20.

Heider, Fritz and Marianne Simmel. 1944. ‘An Experimental Study of Apparent Behaviour’. American Journal of Psychology 57 (2): 243–59.

Heitz, Richard P. 2014. ‘The Speed-Accuracy Tradeoff: History, Physiology, Methodology, and Behavior’. Decision Neuroscience 8: 150. doi:10.3389/fnins.2014.00150.

Helm, Bennett W. 2008. ‘Plural Agents’. Nous 42 (1): 17–49. doi:10.1111/j.1468-0068.2007.00672.x.

Helming, Katharina A., Brent Strickland and Pierre Jacob. 2015. ‘Making Sense of Early False-Belief Understanding’. Trends in Cognitive Sciences. doi:10.1016/j.tics.2014.01.005.

Henmon, V. A. C. 1911. ‘The Relation of the Time of a Judgment to Its Accuracy’. Psychological Review 18 (3): 186–201. doi:10.1037/h0074579.

Hermer, Linda and Elizabeth Spelke. 1996. ‘Modularity and Development: The Case of Spatial Reorientation’. Cognition 61: 195–232.

Hespos, Susan, Gustaf Gredebäck, Claes Von Hofsten and Elizabeth S. Spelke. 2009. ‘Occlusion Is Hard: Comparing Predictive Reaching for Visible and Hidden Objects in Infants and Adults’. Cognitive Science 33 (8): 1483–502. doi:10.1111/j.1551-6709.2009.01051.x.

Heyes, Cecilia M. 2014. ‘False Belief in Infancy: A Fresh Look’. Developmental Science 17 (5): 647–59. doi:10.1111/desc.12148.

Heyes, Cecilia M. and Chris D. Frith. 2014. ‘The Cultural Evolution of Mind Reading’. Science 344 (6190): 1243091. doi:10.1126/science.1243091.

Hoffmann, Almut, Vanessa Rüttler and Andreas Nieder. 2011. ‘Ontogeny of Object Permanence and Object Tracking in the Carrion Crow, Corvus Corone’. Animal Behaviour 82 (2): 359–67. doi:10.1016/j.anbehav.2011.05.012.

Hollebrandse, Bart, Angeliek van Hout and Petra Hendriks. 2012. ‘Children’s First and Second-Order False-Belief Reasoning in a Verbal and a Low-Verbal Task’. Synthese 191 (3): 321–33. doi:10.1007/s11229-012-0169-9.

Hollingworth, Andrew and Steven L. Franconeri. 2009. ‘Object Correspondence Across Brief Occlusion Is Established on the Basis of Both Spatiotemporal and Surface Feature Cues’. Cognition 113 (2): 150–66. doi:10.1016/j.cognition.2009.08.004.

Hood, Bruce, Victoria Cole-Davies and Melanie Dias. 2003. ‘Looking and Search Measures of Object Knowledge in Preschool Children’. Developmental Science 29 (1): 61–70.

Hubbard, Timothy L. 2013. ‘Launching, Entraining, and Representational Momentum: Evidence Consistent with an Impetus Heuristic in Perception of Causality’. Axiomathes 23 (4): 633–43. doi:10.1007/s10516-012-9186-z.

Hughes, Claire and Judy Dunn. 1998. ‘Understanding Mind and Emotion: Longitudinal Associations with Mental-State Talk Between Young Friends’. Developmental Psychology 34 (5): 1026–37. doi:10.1037/0012-1649.34.5.1026.

Hughes, Claire, Keiko K. Fujisawa, Rosie Ensor, Serena Lecce and Rachel Marfleet. 2006. ‘Cooperation and Conversations About the Mind: A Study of Individual Differences in 2-Year-Olds and Their Siblings’. British Journal of Developmental Psychology 24 (1): 53–72.

Hunnius, Sabine and Harold Bekkering. 2010. ‘The Early Development of Object Knowledge: A Study of Infants’ Visual Anticipations During Action Observation’. Developmental Psychology 46 (2): 446–54. doi:10.1037/a0016543.

Hunnius, Sabine and Harold Bekkering. 2014. ‘What Are You Doing? How Active and Observational Experience Shape Infants’ Action Understanding’. Philosophical Transactions of the Royal Society B 369 (1644): 20130490. doi:10.1098/rstb.2013.0490.

Jaakkola, Kelly, Emily Guarino, Mandy Rodriguez, Linda Erb and Marie Trone. 2010. ‘What Do Dolphins (Tursiops Truncatus) Understand About Hidden Objects?’ Animal Cognition 13 (1): 103–20. doi:10.1007/s10071-009-0250-z.

Jackendoff, Ray. 2003a. Foundations of Language: Brain, Meaning, Grammar, Evolution. Oxford: Oxford University Press.

Jackendoff, Ray. 2003b. ‘Précis of Foundations of Language: Brain, Meaning, Grammar, Evolution’. Behavioral and Brain Sciences 26 (6): 651–65. doi:10.1017/S0140525X03000153.

Johnson, Mark H. and John Morton. 1991. Biology and Cognitive Development: The Case of Face Recognition. Oxford: Blackwell.

Johnson, Scott P., J. Gavin Bremner, A. Slater, Uschi Mason, Kirsty Foster and Andrea Cheshire. 2003. ‘Infants’ Perception of Object Trajectories’. Child Development 74 (1): 94–108.

Johnston, Mark. 1992. ‘How to Speak of the Colors’. Philosophical Studies 68 (3): 221–63.

Jonsson, Bert and Claes Von Hofsten. 2003. ‘Infants’ Ability to Track and Reach for Temporarily Occluded Objects’. Developmental Science 6 (1): 86–99. doi:10.1111/1467-7687.00258.

Kahneman, D., Anne Treisman and B. J. Gibbs. 1992. ‘The Reviewing of Object Files: Object-Specific Integration of Information’. Cognitive Psychology 24: 175–219.

Kampis, Dora and Ágnes Melinda Kovács. 2016. ‘Fourteen-Month Old Infants Track Multiple Objects and Object Identity When They Represent Others’ Beliefs’. Paper presented at ICIS, May 26–28, New Orleans.

Kampis, Dora, Eugenio Parise, Gergely Csibra and Ágnes Melinda Kovács. 2015. ‘Neural Signatures for Sustaining Object Representations Attributed to Others in Preverbal Human Infants’. Proceedings of the Royal Society, B 282 (1819): 20151683. doi:10.1098/rspb.2015.1683.

Kanakogi, Yasuhiro and Shoji Itakura. 2011. ‘Developmental Correspondence Between Action Prediction and Motor Ability in Early Infancy’. Nature Communications 2 (June): 341. doi:10.1038/ncomms1342.

Karmiloff-Smith, Annette. 1992. Beyond Modularity: A Developmental Perspective on Cognitive Science. Cambridge, MA: MIT Press.

Kaufman, Jordy, Gergely Csibra and Mark H. Johnson. 2005. ‘Oscillatory Activity in the Infant Brain Reflects Object Maintenance’. Proceedings of the National Academy of Sciences of the United States of America 102 (42): 15271–4. doi:10.1073/pnas.0507626102.

Kawato, Mitsuo. 1999. ‘Internal Models for Motor Control and Trajectory Planning’. Current Opinion in Neurobiology 9 (6): 718–27. doi:10.1016/S0959-4388(99)00028-8.

Keen, Rachel, Neil Berthier, Monica R. Sylvia, Samantha Butler, Patricia K. Prunty and Rachel K. Baker. 2008. ‘Toddlers’ Use of Cues in a Search Task’. Infant and Child Development 17 (3): 249–67. doi:10.1002/icd.550.

Kellman, Philip J. and Elizabeth S. Spelke. 1983. ‘Perception of Partly Occluded Objects in Infancy’. Cognitive Psychology 15 (4): 483–524. doi:10.1016/0010-0285(83)90017-8.

Keren, Gideon and Yaacov Schul. 2009. ‘Two Is Not Always Better Than One’. Perspectives on Psychological Science 4 (6): 533–50. doi:10.1111/j.1745-6924.2009.01164.x.

Kestenbaum, Roberta, Nancy Termine and Elizabeth S. Spelke. 1987. ‘Perception of Objects and Object Boundaries by 3-Month-Old Infants’. British Journal of Developmental Psychology 5 (4): 367–83. doi:10.1111/j.2044-835X.1987.tb01073.x.

Kilner, James M., Yves Paulignan and Sarah-Jayne Blakemore. 2003. ‘An Interference Effect of Observed Biological Movement on Action’. Current Biology 13 (6): 522–5.

Knoblich, Guenther and Natalie Sebanz. 2006. ‘The Social Nature of Perception and Action’. Current Directions in Psychological Science 15 (3): 99–104.

Knoblich, Günther and Natalie Sebanz. 2008. ‘Evolving Intentions for Social Interaction: From Entrainment to Joint Action’. Philosophical Transactions of the Royal Society, B 363: 2021–31.

Knudsen, Birgit and Ulf Liszkowski. 2012a. ‘18-Month-Olds Predict Specific Action Mistakes Through Attribution of False Belief, Not Ignorance, and Intervene Accordingly’. Infancy 17: 672–91. doi:10.1111/j.1532-7078.2011.00105.x.

Knudsen, Birgit and Ulf Liszkowski. 2012b. ‘Eighteen- and 24-Month-Old Infants Correct Others in Anticipation of Action Mistakes’. Developmental Science 15 (1): 113–22. doi:10.1111/j.1467-7687.2011.01098.x.

Kochukhova, Olga and Gustaf Gredebäck. 2010. ‘Preverbal Infants Anticipate That Food Will Be Brought to the Mouth: An Eye Tracking Study of Manual Feeding and Flying Spoons’. Child Development 81 (6): 1729–38. doi:10.1111/j.1467-8624.2010.01506.x.

Koriat, Asher. 1993. ‘How Do We Know That We Know? The Accessibility Model of the Feeling of Knowing’. Psychological Review 100 (4): 609–39. doi:10.1037/0033-295X.100.4.609.

Koriat, Asher. 2007. ‘Metacognition and Consciousness’. In The Cambridge Handbook of Consciousness, edited by P. D. Zelazo, M. Moscovitch and E. Thompson, 289–325. New York: Cambridge University Press. doi:10.1017/CBO9780511816789.012.

Koriat, Asher. 2000. ‘The Feeling of Knowing: Some Metatheoretical Implications for Consciousness and Control’. Consciousness and Cognition 9 (2): 149–71. doi:10.1006/ccog.2000.0433.

Kovács, Ágnes Melinda. 2009. ‘Early Bilingualism Enhances Mechanisms of False-Belief Reasoning’. Developmental Science 12 (1): 48–54. doi:10.1111/j.1467-7687.2008.00742.x.

Kovács, Melinda, Ernő Téglás and Ansgar Denis Endress. 2010. ‘The Social Sense: Susceptibility to Others’ Beliefs in Human Infants and Adults’. Science 330 (6012): 1830–4. doi:10.1126/science.1190792.

Kozhevnikov, Maria and Mary Hegarty. 2001. ‘Impetus Beliefs as Default Heuristics: Dissociation Between Explicit and Implicit Knowledge About Motion’. Psychonomic Bulletin & Review 8 (3): 439–53. doi:10.3758/BF03196179.

Krachun, Carla, Canada Malinda Carpenter, Josep Call and Michael Tomasello. 2010. ‘A New Change-of-Contents False Belief Test: Children and Chimpanzees Compared’. International Journal of Comparative Psychology 23 (2). http://escholarship.org/uc/item/68c0p8dk.

Krushke, J. K. and M. M. Fragassi. 1996. ‘The Perception of Causality: Feature Binding in Interacting Objects’. In Proceedings of the Eighteenth Annual Conference of the Cognitive Science Society, 441–6. Hillsdale, NJ: Erlbaum.

Kulke, Louisa, Josefin Johannsen and Hannes Rakoczy. 2019. ‘Why Can Some Implicit Theory of Mind Tasks Be Replicated and Others Cannot? A Test of Mentalizing Versus Submentalizing Accounts’. PLOS ONE 14 (3): e0213772. doi:10.1371/journal.pone.0213772.

Kulke, Louisa and Hannes Rakoczy. 2018. ‘Implicit Theory of Mind and Overview of Current Replications and Non-Replications’. Data in Brief 16 (Supplement C): 101–4. doi:10.1016/j.dib.2017.11.016.

Kulke, Louisa, Mirjam Reiß, Horst Krist and Hannes Rakoczy. 2017. ‘How Robust Are Anticipatory Looking Measures of Theory of Mind? Replication Attempts Across the Life Span’. Cognitive Development 46: 97–111.

Kulke, Louisa, Britta von Duhn, Dana Schneider and Hannes Rakoczy. 2018. ‘Is Implicit Theory of Mind a Real and Robust Phenomenon? Results from a Systematic Replication Study’. Psychological Science, April, 0956797617747090. doi:10.1177/0956797617747090.

Kundey, Shannon M. A., Andres de Los Reyes, Chelsea Taglang, Ayelet Baruch and Rebecca German. 2010. ‘Domesticated Dogs’ (Canis Familiaris) Use of the Solidity Principle’. Animal Cognition 13 (3): 497–505. doi:10.1007/s10071-009-0300-6.

Kutz, Christopher. 2000. ‘Acting Together’. Philosophy and Phenomenological Research 61 (1): 1–31.

Laurence, Ben. 2011. ‘An Anscombian Approach to Collective Action’. In Essays on Anscombe’s Intention. Cambridge, MA: Harvard University Press.

Leslie, Alan M. 1988. ‘The Necessity of Illusion: Perception and Thought in Infancy’. In Thought Without Language, edited by Lawrence Weiskrantz, 185–210. Oxford: Clarendon Press.

Leslie, Alan M. 2000. ‘Theory of Mind as a Mechanism of Selective Attention’. In The New Cognitive Neurosciences, edited by Michael S. Gazzaniga, 1235–47. Cambridge, MA: MIT Press.

Leslie, Alan M., Tim P. German and Pamela Polizzi. 2005. ‘Belief-Desire Reasoning as a Process of Selection’. Cognitive Psychology 50: 45–85.

Leslie, Alan M. and Pamela Polizzi. 1998. ‘Inhibitory Processing in the False Belief Task: Two Conjectures’. Developmental Science 1 (2): 247–53.

Leslie, Alan M., Fei Xu, Patrice D. Tremoulet and Brian J. Scholl. 1998. ‘Indexing and the Object Concept: Developing “What” and “Where” Systems’. Trends in Cognitive Sciences 2 (1).

Lidz, Jeffrey and Sandra Waxman. 2004. ‘Reaffirming the Poverty of the Stimulus Argument: A Reply to the Replies’. Cognition 93 (2): 157–65. doi:10.1016/j.cognition.2004.02.001.

Lidz, Jeffrey, Sandra Waxman and Jennifer Freedman. 2003. ‘What Infants Know About Syntax but Couldn’t Have Learned: Experimental Evidence for Syntactic Structure at 18 Months’. Cognition 89 (3): 295–303. doi:10.1016/S0010-0277(03)00116-1.

Linnebo, Øystein. 2005. ‘Plural Quantification’. In The Stanford Encyclopedia of Philosophy (Spring 2005 Edition), edited by Edward N. Zalta. Stanford, CA: Metaphysics Research Lab, Stanford University.

Lohmann, Heidemarie and Michael Tomasello. 2003. ‘The Role of Language in the Development of False Belief Understanding: A Training Study’. Child Development 74 (4): 1130–44. doi:10.1111/1467-8624.00597.

Low, Jason. 2010. ‘Preschoolers’ Implicit and Explicit False-Belief Understanding: Relations with Complex Syntactical Mastery’. Child Development 81 (2): 597–615. doi:10.1111/j.1467-8624.2009.01418.x.

Low, Jason, Ian A. Apperly, Stephen A. Butterfill and Hannes Rakoczy. 2016. ‘Cognitive Architecture of Belief Reasoning in Children and Adults: A Primer on the Two-Systems Account’. Child Development Perspectives 10 (3): 184–9. doi:10.1111/cdep.12183.

Low, Jason, William Drummond, Andrew Walmsley and Bo Wang. 2014. ‘Representing How Rabbits Quack and Competitors Act: Limits on Preschoolers’ Efficient Ability to Track Perspective’. Child Development 85 (4).

Low, Jason and Joseph Watts. 2013. ‘Attributing False-Beliefs About Object Identity Is a Signature Blindspot in Humans’ Efficient Mindreading System’. Psychological Science 24 (3): 305–11.

Ludwig, Kirk. 2015. ‘Shared Agency in Modest Sociality’. Journal of Social Ontology 1 (1): 7–15.

Luo, Yuyan. 2011. ‘Three-Month-Old Infants Attribute Goals to a Non-Human Agent’. Developmental Science 14 (2): 453–60. doi:10.1111/j.1467-7687.2010.00995.x.

Mameli, Matteo and Patrick Bateson. 2011. ‘An Evaluation of the Concept of Innateness’. Philosophical Transactions of the Royal Society, B: Biological Sciences 366 (1563): 436–43. doi:10.1098/rstb.2010.0174.

Mandler, Jean M. 1992. ‘How to Build a Baby: II. Conceptual Primitives’. Psychological Review 99 (4): 587–604.

Marr, David. 1982. Vision: A Computational Investigation into the Human Representation and Processing of Visual Information. San Francisco: W.H. Freeman.

Mash, Clay, Elizabeth Novak, Neil E. Berthier and Rachel Keen. 2006. ‘What Do Two-Year-Olds Understand About Hidden-Object Events?’ Developmental Psychology 42 (2): 263–71. doi:10.1037/0012-1649.42.2.263.

Matthews, R. J. 2007. The Measure of Mind: Propositional Attitudes and Their Attribution. Oxford: Oxford University Press.

McCloskey, Michael, Allyson Washburn and Linda Felch. 1983. ‘Intuitive Physics: The Straight-down Belief and Its Origin’. Journal of Experimental Psychology: Learning, Memory, and Cognition 9 (4): 636–49. doi:10.1037/0278-7393.9.4.636.

McCurry, Sarah, Teresa Wilcox and Rebecca Woods. 2009. ‘Beyond the Search Barrier: A New Task for Assessing Object Individuation in Young Infants’. Infant Behavior and Development 32 (4): 429–36. doi:10.1016/j.infbeh.2009.07.002.

McGeer, Victoria. 1996. ‘Is “Self-Knowledge” an Empirical Problem? Renegotiating the Space of Philosophical Explanation’. The Journal of Philosophy 93 (10): 483–515.

Meltzoff, Andrew N. 2007. ‘“Like Me”: A Foundation for Social Cognition’. Developmental Science 10 (1): 126–34.

Meltzoff, Andrew N. and M. Keith Moore. 1998. ‘Object Representation, Identity, and the Paradox of Early Permanence: Steps Toward a New Framework’. Infant Behavior and Development 21 (2): 201–35.

Melzer, Anne, Wolfgang Prinz and Moritz M. Daum. 2012. ‘Production and Perception of Contralateral Reaching: A Close Link by 12 Months of Age’. Infant Behavior and Development 35 (3): 570–9. doi:10.1016/j.infbeh.2012.05.003.

Meristo, Marek, Kerstin W. Falkman, Erland Hjelmquist, Mariantonia Tedoldi, Luca Surian and Michael Siegal. 2007. ‘Language Access and Theory of Mind Reasoning: Evidence from Deaf Children in Bilingual and Oralist Environments’. Developmental Psychology 43 (5): 1156–69. doi:10.1037/0012-1649.43.5.1156.

Meristo, Marek, Gary Morgan, Alessandra Geraci, Laura Iozzi, Erland Hjelmquist, Luca Surian and Michael Siegal. 2012. ‘Belief Attribution in Deaf and Hearing Infants’. Developmental Science 15 (5): 633–40. doi:10.1111/j.1467-7687.2012.01155.x.

Meristo, Marek, Karin Strid and Erland Hjelmquist. 2016. ‘Early Conversational Environment Enables Spontaneous Belief Attribution in Deaf Children’. Cognition 157: 139–45. doi:10.1016/j.cognition.2016.08.023.

Meyer, Marlene, Robrecht P. R. D. van der Wel and Sabine Hunnius. 2016. ‘Planning My Actions to Accommodate Yours: Joint Action Development During Early Childhood’. Philosophical Transactions of the Royal Society, B 371 (1693): 20150371. doi:10.1098/rstb.2015.0371.

Milligan, Karen, Janet Wilde Astington and Lisa Ain Dack. 2007. ‘Language and Theory of Mind: Meta-Analysis of the Relation Between Language Ability and False-Belief Understanding’. Child Development 78 (2): 622–46. doi:10.1111/j.1467-8624.2007.01018.x.

Mitroff, Stephen R. and George A. Alvarez. 2007. ‘Space and Time, Not Surface Features, Guide Object Persistence’. Psychonomic Bulletin & Review 14 (6): 1199–204. doi:10.3758/BF03193113.

Mitroff, Stephen R., Jason T. Arita and Mathias S. Fleck. 2009. ‘Staying in Bounds: Contextual Constraints on Object-File Coherence’. Visual Cognition 17 (1–2): 195–211. doi:10.1080/13506280802103457.

Mitroff, Stephen R., Brian J. Scholl and Karen Wynn. 2005. ‘The Relationship Between Object Files and Conscious Perception’. Cognition 96 (1): 67–92. doi:10.1016/j.cognition.2004.03.008.

Moeller, Mary Pat and Brenda Schick. 2006. ‘Relations Between Maternal Input and Theory of Mind Understanding in Deaf Children’. Child Development 77 (3): 751–66. doi:10.1111/j.1467-8624.2006.00901.x.

Moll, Henrike and Michael Tomasello. 2007. ‘Cooperation and Human Cognition: The Vygotskian Intelligence Hypothesis’. Philosophical Transactions of the Royal Society, B 362 (1480): 639–48.

Moore, Cathleen M., Teresa Stephens and Elisabeth Hein. 2010. ‘Features, as Well as Space and Time, Guide Object Persistence’. Psychonomic Bulletin & Review 17 (5): 731–6. doi:10.3758/PBR.17.5.731.

Moore, M. Keith and Andrew N. Meltzoff. 2008. ‘Factors Affecting Infants’ Manual Search for Occluded Objects and the Genesis of Object Permanence’. Infant Behavior and Development 31 (2): 168–80. doi:10.1016/j.infbeh.2007.10.006.

Morgan, Gary, Marek Meristo, Wolfgang Mann, Erland Hjelmquist, Luca Surian and Michael Siegal. 2014. ‘Mental State Language and Quality of Conversational Experience in Deaf and Hearing Children’. Cognitive Development 29 (January): 41–9. doi:10.1016/j.cogdev.2013.10.002.

Munakata, Yuko. 2001. ‘Graded Representations in Behavioral Dissociations’. Trends in Cognitive Sciences 5 (7): 309–15.

Munakata, Yuko, James L. McClelland, Mark H. Johnson and Robert S. Siegler. 1997. ‘Rethinking Infant Knowledge: Toward an Adaptive Process Account of Successes and Failures in Object Permanence Tasks’. Psychological Review 104 (4): 686–713. doi:10.1037/0033-295X.104.4.686.

Nagel, Jennifer, Raymond Mar and Valerie San Juan. 2013. ‘Authentic Gettier Cases: A Reply to Starmans and Friedman’. Cognition 129 (3): 666–9. doi:10.1016/j.cognition.2013.08.016.

Naughtin, Claire K., Kristina Horne, Dana Schneider, Dustin Venini, Ashley York and Paul E. Dux. 2017. ‘Do Implicit and Explicit Belief Processing Share Neural Substrates?’ Human Brain Mapping 38 (9): 4760–72. doi:10.1002/hbm.23700.

Needham, Amy. 1998. ‘Infants’ Use of Featural Information in the Segregation of Stationary Objects’. Infant Behavior and Development 21 (1): 47–76. doi:10.1016/S0163-6383(98)90054-6.

Needham, Amy. 1999. ‘The Role of Shape in 4-Month-Old Infants’ Object Segregation’. Infant Behavior and Development 22 (2): 161–78. doi:10.1016/S0163-6383(99)00008-9.

Needham, Amy, Tracy Barrett and Karen Peterman. 2002. ‘A Pick-Me-up for Infants’ Exploratory Skills: Early Simulated Experiences Reaching for Objects Using “Sticky Mittens” Enhances Young Infants’ Object Exploration Skills’. Infant Behavior and Development 25 (3): 279–95. doi:10.1016/S0163-6383(02)00097-8.

Noles, Nicholaus S., Brian J. Scholl and Stephen R. Mitroff. 2005. ‘The Persistence of Object File Representations’. Perception & Psychophysics 67 (2): 324–34. doi:10.3758/BF03206495.

Oberle, Crystal D., Michael K. McBeath, Sean C. Madigan and Thomas G. Sugar. 2005. ‘The Galileo Bias: A Naive Conceptual Belief That Influences People’s Perceptions and Performance in a Ball-Dropping Task’. Journal of Experimental Psychology: Learning, Memory, and Cognition 31 (4): 643–53. doi:10.1037/0278-7393.31.4.643.

Odic, Darko, Oliver Roth and Jonathan I. Flombaum. 2012. ‘The Relationship Between Apparent Motion and Object Files’. Visual Cognition 20 (9): 1052–81. doi:10.1080/13506285.2012.721405.

Oktay-Gür, Nese, Alexandra Schulz and Hannes Rakoczy. 2018. ‘Children Exhibit Different Performance Patterns in Explicit and Implicit Theory of Mind Tasks’. Cognition 173: 60–74. doi:10.1016/j.cognition.2018.01.001.

Onishi, Kristine H. and Renée Baillargeon. 2005. ‘Do 15-Month-Old Infants Understand False Beliefs?’ Science 308 (8): 255–8.

Pacherie, Elisabeth. 2013. ‘Intentional Joint Agency: Shared Intention Lite’. Synthese 190 (10): 1817–39. doi:10.1007/s11229-013-0263-7.

Palmer, Stephen E. 1999. Vision Science: Photons to Phenomenology. Cambridge, MA: MIT Press.

Pattison, Kristina F., Holly C. Miller, Rebecca Rayburn-Reeves and Thomas Zentall. 2010. ‘The Case of the Disappearing Bone: Dogs’ Understanding of the Physical Properties of Objects’. Behavioural Processes, Comparative Cognition: A Tribute to the Contributions of Donald A. Riley, 85 (3): 278–82. doi:10.1016/j.beproc.2010.06.016.

Paulus, Markus. 2016. ‘The Development of Action Planning in a Joint Action Context’. Developmental Psychology 52 (7): 1052–63. doi:10.1037/dev0000139.

Paulus, Markus, Sabine Hunnius, Carolien van Wijngaarden, Sven Vrins, Iris van Rooij and Harold Bekkering. 2011. ‘The Role of Frequency Information and Teleological Reasoning in Infants’ and Adults’ Action Prediction’. Developmental Psychology 47 (4): 976–83. doi:10.1037/a0023785.

Perner, Josef. 1991. The Representational Mind. Brighton: Harvester.

Perner, Josef and Johannes Roessler. 2012. ‘From Infants’ to Children’s Appreciation of Belief’. Trends in Cognitive Sciences 16 (10): 519–25. doi:10.1016/j.tics.2012.08.004.

Peterson, Candida C. 2009. ‘Development of Social-Cognitive and Communication Skills in Children Born Deaf’. Scandinavian Journal of Psychology 50 (5): 475–83. doi:10.1111/j.1467-9450.2009.00750.x.

Peterson, Candida C. and Michael Siegal. 2000. ‘Insights into Theory of Mind from Deafness and Autism’. Mind & Language 15 (1): 123–45. doi:10.1111/1468-0017.00126.

Peterson, Candida and Virginia Slaughter. 2003. ‘Opening Windows into the Mind: Mothers’ Preferences for Mental State Explanations and Children’s Theory of Mind’. Cognitive Development 18 (3): 399–429. doi:10.1016/S0885-2014(03)00041-8.

Phillips, Jonathan, Desmond C. Ong, Andrew D. R. Surtees, Yijing Xin, Samantha Williams, Rebecca Saxe and Michael C. Frank. 2015. ‘A Second Look at Automatic Theory of Mind: Reconsidering Kovács, Téglás, and Endress (2010)’. Psychological Science 26 (9): 1353–67. doi:10.1177/0956797614558717.

Polak, Alan and Paul Harris. 1999. ‘Deception by Young Children Following Noncompliance’. Developmental Psychology 35 (2): 561–8.

Poulin-Dubois, Diane, Hannes Rakoczy, Kimberly Burnside, Cristina Crivello, Sebastian Dörrenberg, Katheryn Edwards, Horst Krist, et al. 2018. ‘Do Infants Understand False Beliefs? We Don’t Know Yet. A Commentary on Baillargeon, Buttelmann and Southgate’s Commentary’. Cognitive Development 48 (October): 302–15. doi:10.1016/j.cogdev.2018.09.005.

Poulin-Dubois, Diane and Jessica Yott. 2017. ‘Probing the Depth of Infants’ Theory of Mind: Disunity in Performance Across Paradigms’. Developmental Science doi:10.1111/desc.12600.

Powell, Lindsey J., Kathryn Hobbs, Alexandros Bardis, Susan Carey and Rebecca Saxe. 2017. ‘Replications of Implicit Theory of Mind Tasks with Varying Representational Demands’. Cognitive Development, October. doi:10.1016/j.cogdev.2017.10.004.

Premack, David. 1990. ‘The Infant’s Theory of Self-Propelled Objects’. Cognition 36 (1): 1–16.

Premack, David and Guy Woodruff. 1978. ‘Does the Chimpanzee Have a Theory of Mind?’ Behavioral and Brain Sciences 1 (04): 515–26. doi:10.1017/S0140525X00076512.

Pullum, Geoffrey K. and Barbara C. Scholz. 2002. ‘Empirical Assessment of Stimulus Poverty Arguments’. The Linguistic Review 18 (1–2). doi:10.1515/tlir.19.1-2.9.

Pyers, Jennie E. and Ann Senghas. 2009. ‘Language Promotes False-Belief Understanding Evidence from Learners of a New Sign Language’. Psychological Science 20 (7): 805–12. doi:10.1111/j.1467-9280.2009.02377.x.

Pylyshyn, Zenon W. and Ron W. Storm. 1988. ‘Tracking Multiple Independent Targets: Evidence for a Parallel Tracking Mechanism’. Spatial Vision 3 (3): 179–97.

Qureshi, Adam, Ian A. Apperly and Dana Samson. 2010. ‘Executive Function Is Necessary for Perspective Selection, Not Level-1 Visual Perspective Calculation: Evidence from a Dual-Task Study of Adults’. Cognition 117 (2): 230–6.

Rakoczy, Hannes. 2006. ‘Pretend Play and the Development of Collective Intentionality’. Cognitive Systems Research 7 (2–3): 113–27. doi:10.1016/j.cogsys.2005.11.008.

Rakoczy, Hannes. 2010. ‘Executive Function and the Development of Belief-Desire Psychology’. Developmental Science 13 (4): 648–61. doi:10.1111/j.1467-7687.2009.00922.x.

Rakoczy, Hannes, Felix Warneken and Michael Tomasello. 2007. ‘“This Way!”, “No! That Way!”—3-Year Olds Know That Two People Can Have Mutually Incompatible Desires’. Cognitive Development 22 (1): 47–68. doi:10.1016/j.cogdev.2006.08.002.

Ramenzoni, Verónica C., Tehran J. Davis, Michael A. Riley, Kevin Shockley and Aimee A. Baker. 2011. ‘Joint Action in a Cooperative Precision Task: Nested Processes of Intrapersonal and Interpersonal Coordination’. Experimental Brain Research 211 (3–4): 447–57. doi:10.1007/s00221-011-2653-8.

Razza, Rachel A. and Clancy Blair. 2009. ‘Associations Among False-Belief Understanding, Executive Function, and Social Competence: A Longitudinal Analysis’. Journal of Applied Developmental Psychology 30 (3): 332–43. doi:10.1016/j.appdev.2008.12.020.

Reid, Thomas. 1785a. An Inquiry into the Human Mind. 4th edn. London: T. Cadell et al.

Reid, Thomas. 1785b. Essays on the Intellectual Powers of Man. Edinburgh: John Bell & G. Robinson.

Reisenzein, Rainer. 2000. ‘The Subjective Experience of Surprise’. In The Message Within: The Role of Subjective Experience in Social Cognition and Behavior, edited by H. Bless and J. P. Forgas, 262–79. Hove: Psychology Press.

Richardson, Daniel C. and Natasha Z. Kirkham. 2004. ‘Multimodal Events and Moving Locations: Eye Movements of Adults and 6-Month-Olds Reveal Dynamic Spatial Indexing’. Journal of Experimental Psychology: General 133 (1): 46–62. doi:10.1037/0096-3445.133.1.46.

Riggs, K. J. and E. J. Robinson. 1995. ‘What People Say and What They Think: Children’s Judgements of False Belief in Relation to Their Recall of False Messages’. British Journal of Developmental Psychology 13 (3): 271–84. doi:10.1111/j.2044-835X.1995.tb00679.x.

Rips, Lance J. 2011. ‘Causation from Perception’. Perspectives on Psychological Science 6 (1): 77–97. doi:10.1177/1745691610393525.

Rizzolatti, Giacomo and Corrado Sinigaglia. 2008. Mirrors in the Brain: How Our Minds Share Actions, Emotions. Oxford: Oxford University Press.

Rizzolatti, Giacomo and Corrado Sinigaglia. 2016. ‘The Mirror Mechanism: A Basic Principle of Brain Function’. Nature Reviews Neuroscience, advance online publication. doi:10.1038/nrn.2016.135.

Rochat, Philippe, Tricia Striano and R. Morgan. 2004. ‘Who Is Doing What to Whom? Young Infants’ Developing Sense of Social Causality in Animated Displays’. Perception 33 (3): 355–69. doi:10.1068/p3389.

Roessler, Johannes and Josef Perner. 2013. ‘Teleology: Belief as Perspective’. In Understanding Other Minds: Perspectives from Developmental Social Neuroscience. 35–50. Oxford: Oxford University Press.

Rosander, Kerstin and Claes von Hofsten. 2004. ‘Infants’ Emerging Ability to Represent Occluded Object Motion’. Cognition 91 (1): 1–22. doi:10.1016/S0010-0277(03)00166-5.

Rosenbaum, David A. 2010. Human Motor Control. 2nd edn. San Diego, CA: Academic Press.

Rubio-Fernández, Paula. 2013. ‘Perspective Tracking in Progress: Do Not Disturb’. Cognition 129 (2): 264–72. doi:10.1016/j.cognition.2013.07.005.

Rubio-Fernández, Paula and Bart Geurts. 2012. ‘How to Pass the False-Belief Task Before Your Fourth Birthday’. Psychological Science, November, 0956797612447819. doi:10.1177/0956797612447819.

Rubio-Fernández, Paula and Bart Geurts. 2016. ‘Don’t Mention the Marble! The Role of Attentional Processes in False-Belief Tasks’. Review of Philosophy and Psychology 7 (4): 835–50. doi:10.1007/s13164-015-0290-z.

Ruffman, Ted. 2014. ‘To Belief or Not Belief: Children’s Theory of Mind’. Developmental Review 34 (3): 265–93. doi:10.1016/j.dr.2014.04.001.

Ruffman, Ted, Wendy Garnham, Arlina Import and Dan Connolly. 2001. ‘Does Eye Gaze Indicate Implicit Knowledge of False Belief? Charting Transitions in Knowledge’. Journal of Experimental Child Psychology 80: 201–24.

Ruffman, Ted, Josef Perner, Mika Naito, Lindsay Parkin and Wendy A. Clements. 1998. ‘Older (but Not Younger) Siblings Facilitate False Belief Understanding’. Developmental Psychology 34 (1): 161–74. doi:10.1037/0012-1649.34.1.161.

Ruffman, Ted, Lance Slade and Elena Crowe. 2002. ‘The Relation Between Children’s and Mothers’ Mental State Language and Theory-of-Mind Understanding’. Child Development 73 (3): 734–51. doi:10.1111/1467-8624.00435.

Sabbagh, Mark A., Fen Xu, Stephanie M. Carlson, Louis J. Moses and Kang Lee. 2006. ‘The Development of Executive Functioning and Theory of Mind: A Comparison of Chinese and U.S. Preschoolers’. Psychological Science 17 (1): 74–81. doi:10.1111/j.1467-9280.2005.01667.x.

Samuels, Richard. 2004. ‘Innateness in Cognitive Science’. Trends in Cognitive Sciences 8 (3): 136–41.

San Juan, Valerie and Janet Wilde Astington. 2012. ‘Bridging the Gap Between Implicit and Explicit Understanding: How Language Development Promotes the Processing and Representation of False Belief’. British Journal of Developmental Psychology 30 (1): 105–22. doi:10.1111/j.2044-835X.2011.02051.x.

Santos, Laurie R and Bruce M Hood. 2009. ‘Object Representation as a Central Issue in Cognitive Science’. In The Origins of Object Knowledge, edited by Bruce M. Hood and Laurie R. Santos, 1–23. Oxford: Oxford University Press.

Santos, Laurie R., David Seelig and Marc D. Hauser. 2006. ‘Cotton-Top Tamarins’ (Saguinus Oedipus) Expectations About Occluded Objects: A Dissociation Between Looking and Reaching Tasks’. Infancy 9 (2): 147–71. doi:10.1207/s15327078in0902_4.

Saxe, Rebecca, Tania Tzelnic and Susan Carey. 2006. ‘Five-Month-Old Infants Know Humans Are Solid, Like Inanimate Objects’. Cognition 101 (1): B1–B8. doi:10.1016/j.cognition.2005.10.005.

Schick, Brenda, Peter de Villiers, Jill de Villiers and Robert Hoffmeister. 2007. ‘Language and Theory of Mind: A Study of Deaf Children’. Child Development 78 (2): 376–96. doi:10.1111/j.1467-8624.2007.01004.x.

Schlottmann, Anne and Elizabeth Ray. 2010. ‘Goal Attribution to Schematic Animals: Do 6-Month-Olds Perceive Biological Motion as Animate?’ Developmental Science 13 (1): 1–10. doi:10.1111/j.1467-7687.2009.00854.x.

Schmidt, Richard C. and Michael J. Richardson. 2008. ‘Dynamics of Interpersonal Coordination’. In Coordination: Neural, Behavioral and Social Dynamics, edited by Armin Fuchs and Viktor K. Jirsa, 280–308. Berlin: Springer. www.springerlink.com/content/627nx655445v6q82/.

Schneider, Dana, Rebecca Lam, Andrew P. Bayliss and Paul E. Dux. 2012. ‘Cognitive Load Disrupts Implicit Theory-of-Mind Processing’. Psychological Science 23 (8): 842–7. doi:10.1177/0956797612439070.

Schneider, Dana, Zoie E. Nott and Paul E. Dux. 2014. ‘Task Instructions and Implicit Theory of Mind’. Cognition 133 (1): 43–7. doi:10.1016/j.cognition.2014.05.016.

Schneider, Dana, Virginia P. Slaughter and Paul E. Dux. 2017. ‘Current Evidence for Automatic Theory of Mind Processing in Adults’. Cognition 162 (May): 27–31. doi:10.1016/j.cognition.2017.01.018.

Schneider, D., A. P. Bayliss, S. I. Becker and P. E. Dux. 2012. ‘Eye Movements Reveal Sustained Implicit Processing of Others’ Mental States’. Journal of Experimental Psychology: General 141 (3): 433–8.

Scholl, Brian J. 2005. ‘Innateness and (Bayesian) Visual Perception’. In The Innate Mind: Structure and Contents, edited by Peter Carruthers, Stephen Laurence and Stephen Stich, 34–52. Oxford: Oxford University Press.

Scholl, Brian J. 2007. ‘Object Persistence in Philosophy and Psychology’. Mind & Language 22 (5): 563–91. doi:10.1111/j.1468-0017.2007.00321.x.

Scholl, Brian J. and J. I. Flombaum. 2010. ‘Object Persistence’. In Encyclopedia of Perception, edited by B. Goldstein, 2: 653–7. Thousand Oaks, CA: Sage.

Scholl, Brian J. and Tao Gao. 2013. ‘Perceiving Animacy and Intentionality: Visual Processing or Higher-Level Judgment?’ In Social Perception: Detection and Interpretation of Animacy, Agency, and Intention, edited by M. D. Rutherford and V. A. Kuhlmeier, 197–230. Cambridge, MA: MIT Press.

Scholl, Brian J. and Alan M. Leslie. 1999. ‘Explaining the Infant’s Object Concept: Beyond the Perception/Cognition Dichotomy’. In What Is Cognitive Science?, edited by E. LePore and Zenon W. Pylyshyn, 26–73. Oxford: Blackwell.

Scholl, Brian J. and Zenon W. Pylyshyn. 1999. ‘Tracking Multiple Items Through Occlusion: Clues to Visual Objecthood’. Cognitive Psychology 38 (2): 259–90. http://cat.inist.fr/?aModele=afficheN\&cpsidt=1723632.

Scholl, Brian J. and Patrice D. Tremoulet. 2000. ‘Perceptual Causality and Animacy’. Trends in Cognitive Sciences 4 (8): 299–309.

Schöner, Gregor and Evelina Dineva. 2007. ‘Dynamic Instabilities as Mechanisms for Emergence’. Developmental Science 10 (1): 69–74. doi:10.1111/j.1467-7687.2007.00566.x.

Schöner, Gregor and Esther Thelen. 2006. ‘Using Dynamic Field Theory to Rethink Infant Habituation’. Psychological Review 113 (2): 273–99. doi:10.1037/0033-295X.113.2.273.

Scott, R. and Renée Baillargeon. 2009. ‘Which Penguin Is This? Attributing False Beliefs About Object Identity at 18 Months’. Child Development 80 (4): 1172–96.

Scott, Rose M. 2017. ‘Surprise! 20-Month-Old Infants Understand the Emotional Consequences of False Beliefs’. Cognition 159 (February): 33–47. doi:10.1016/j.cognition.2016.11.005.

Scott, Rose M., Renée Baillargeon, Hyun-joo Song and Alan M. Leslie. 2010. ‘Attributing False Beliefs About Non-Obvious Properties at 18 Months’. Cognitive Psychology 61 (4): 366–95. doi:10.1016/j.cogpsych.2010.09.001.

Scott, Rose M., Zijing He, Renée Baillargeon and Denise Cummins.. 2012. ‘False-Belief Understanding in 2.5-Year-Olds: Evidence from Two Novel Verbal Spontaneous-Response Tasks’. Developmental Science 15 (2): 181–93. doi:10.1111/j.1467-7687.2011.01103.x, 10.1111/j.1467-7687.2011.01103.x.

Scott, Rose M., Joshua C. Richman and Renée Baillargeon. 2015. ‘Infants Understand Deceptive Intentions to Implant False Beliefs About Identity: New Evidence for Early Mentalistic Reasoning’. Cognitive Psychology 82 (November): 32–56. doi:10.1016/j.cogpsych.2015.08.003.

Scott, Ryan B. and Zoltán Dienes. 2008. ‘The Conscious, the Unconscious, and Familiarity’. Journal of Experimental Psychology: Learning, Memory, and Cognition 34 (5): 1264–88.

Searle, John R. 1990. ‘Collective Intentions and Actions’. In Intentions in Communication, edited by P. Cohen, J. Morgan and M. E. Pollack, 90–105. Cambridge: Cambridge University Press.

Sherman, Jeffrey W., Bertram Gawronski and Yaacov Trope, eds. 2014. Dual-Process Theories of the Social Mind. New York: Guilford Press.

Shinskey, Jeanne L. 2012. ‘Disappearing Décalage: Object Search in Light and Dark at 6 Months’. Infancy 17 (3): 272–94. doi:10.1111/j.1532-7078.2011.00078.x.

Shinskey, Jeanne and Yuko Munakata. 2001. ‘Detecting Transparent Barriers: Clear Evidence Against the Means-End Deficit Account of Search Failures’. Infancy 2 (3): 395–404.

Sidarus, Nura, Valérian Chambon and Patrick Haggard. 2013. ‘Priming of Actions Increases Sense of Control over Unexpected Outcomes’. Consciousness and Cognition 22 (4): 1403–11. doi:10.1016/j.concog.2013.09.008.

Sidarus, Nura, Matti Vuorre and Patrick Haggard. 2017. ‘How Action Selection Influences the Sense of Agency: An ERP Study’. NeuroImage 150 (April): 1–13. doi:10.1016/j.neuroimage.2017.02.015.

Siegal, Michael and Karen Beattie. 1991. ‘Where to Look First for Children’s Knowledge of False Beliefs’. Cognition 38 (1): 1–12. doi:10.1016/0010-0277(91)90020-5.

Sinigaglia, Corrado and Stephen A. Butterfill. 2016. ‘Motor Representation in Goal Ascription’. In Foundations of Embodied Cognition 2: Conceptual and Interactive Embodiment, edited by Yann Coello and Martin H. Fischer, 149–64. Hove: Psychology Press.

Sinigaglia, Corrado and Laura Sparaci. 2008. ‘The Mirror Roots of Social Cognition’. Acta Philosophica 17 (2): 307–30.

Sirois, Sylvain and Iain R. Jackson. 2012. ‘Pupil Dilation and Object Permanence in Infants’. Infancy 17 (1): 61–78. doi:10.1111/j.1532-7078.2011.00096.x.

Sirois, Sylvain and Denis Mareschal. 2002. ‘Models of Habituation in Infancy’. Trends in Cognitive Sciences 6 (7): 293–8. doi:10.1016/S1364-6613(02)01926-5.

Skerry, Amy E., Susan E. Carey and Elizabeth S. Spelke. 2013. ‘First-Person Action Experience Reveals Sensitivity to Action Efficiency in Prereaching Infants’. Proceedings of the National Academy of Sciences 110 (46): 18728–33. doi:10.1073/pnas.1312322110.

Slaughter, Virginia and Alison Gopnik. 1996. ‘Conceptual Coherence in the Child’s Theory of Mind: Training Children to Understand Belief’. Child Development 67: 2967–88.

Smith, A. D. 2001. ‘Perception and Belief’. Philosophy and Phenomenological Research 62 (2): 283–309.

Smith, Joel. 2010. ‘Seeing Other People’. Philosophy and Phenomenological Research 81 (3): 731–48. doi:10.1111/j.1933-1592.2010.00392.x.

Smith, Linda B. 2005. ‘Cognition as a Dynamic System: Principles from Embodiment’. Developmental Review, Development as Self-organization: New Approaches to the Psychology and Neurobiology of Development, 25 (34): 278–98. doi:10.1016/j.dr.2005.11.001.

Smith, Thomas H. 2015. ‘Shared Agency on Gilbert and Deep Continuity’. Journal of Social Ontology 1 (1): 49–57.

Sommerville, Jessica A., Elina A. Hildebrand and Catharyn C. Crane. 2008. ‘Experience Matters: The Impact of Doing Versus Watching on Infants’ Subsequent Perception of Tool-Use Events’. Developmental Psychology 44 (5): 1249–56. doi:10.1037/a0012296.

Sommerville, Jessica A., Amanda L. Woodward and Amy Needham. 2005. ‘Action Experience Alters 3-Month-Old Infants’ Perception of Others’ Actions’. Cognition 96 (1): B1–B11. doi:16/j.cognition.2004.07.004.

Southgate, Victoria. 2013. ‘Do Infants Provide Evidence That the Mirror System Is Involved in Action Understanding?’ Consciousness and Cognition 22 (3): 1114–21. doi:10.1016/j.concog.2013.04.008.

Southgate, Victoria, Coralie Chevallier and Gergely Csibra. 2010. ‘Seventeen-Month-Olds Appeal to False Beliefs to Interpret Others’ Referential Communication’. Developmental Science 13 (6): 907–12. doi:10.1111/j.1467-7687.2009.00946.x.

Southgate, Victoria, A. Senju and Csibra Gergely. 2007. ‘Action Anticipation Through Attribution of False Belief by Two-Year-Olds’. Psychological Science 18 (7): 587–92.

Southgate, Victoria and Angelina Vernetti. 2014. ‘Belief-Based Action Prediction in Preverbal Infants’. Cognition 130 (1): 1–10. doi:10.1016/j.cognition.2013.08.008.

Spelke, Elizabeth. 1988. ‘Where Perceiving Ends and Thinking Begins: The Apprehension of Objects in Infancy’. In Perceptual Development in Early Infancy, edited by A. Yonas, 197–234. Hillsdale, NJ: Erlbaum.

Spelke, Elizabeth. 1990. ‘Principles of Object Perception’. Cognitive Science 14: 29–56.

Spelke, Elizabeth. 1994. ‘Initial Knowledge: Six Suggestions’. Cognition 50 (1–3): 431–45.

Spelke, Elizabeth. 1998. ‘Nativism, Empiricism, and the Origins of Knowledge’. Infant Behavior and Development 21 (2): 181–200.

Spelke, Elizabeth. 2000. ‘Core Knowledge’. American Psychologist 55: 1233–43.

Spelke, Elizabeth. 2003. ‘What Makes Us Smart?’ In Advances in the Study of Language and Thought, edited by D. Gentner and S. Goldin-Meadow. Cambridge, MA: MIT Press.

Spelke, Elizabeth S., Karen Breinlinger, Janet Macomber and Kristen Jacobson. 1992. ‘Origins of Knowledge’. Psychological Review 99 (4): 605–32. doi:10.1037/0033-295X.99.4.605.

Spelke, Elizabeth and Susan Hespos. 2001. ‘Continuity, Competence, and the Object Concept’. In Language, Brain, and Cognitive Development, edited by Emmanuel Dupoux. Cambridge, MA: MIT Press.

Spelke, Elizabeth S., Roberta Kestenbaum, Daniel J. Simons and Debra Wein. 1995. ‘Spatiotemporal Continuity, Smoothness of Motion and Object Identity in Infancy’. British Journal of Developmental Psychology 13 (2): 113–42. doi:10.1111/j.2044-835X.1995.tb00669.x.

Spelke, Elizabeth S., C. von Hofsten and Roberta Kestenbaum. 1989. ‘Object Perception and Object-Directed Reaching in Infancy: Interaction of Spatial and Kinetic Information for Object Boundaries’. Developmental Psychology 25: 185–96. Available at : www.wjh.harvard.edu/~lds/pdfs/Spelke\%20von\%20Hofsten\%20Kestenbaum\%201989.pdf.

Sperber, Dan. 1997. ‘Intuitive and Reflective Beliefs’. Mind and Language 12 (1): 67–83.

Starmans, Christina and Ori Friedman. 2012. ‘The Folk Conception of Knowledge’. Cognition 124 (3): 272–83. doi:10.1016/j.cognition.2012.05.017.

Starmans, Christina and Ori Friedman. 2013. ‘Taking “Know” for an Answer: A Reply to Nagel, San Juan, and Mar’. Cognition 129 (3): 662–5. doi:10.1016/j.cognition.2013.05.009.

Stich, Stephen. 1978. ‘Beliefs and Subdoxastic States’. Philosophy of Science 45: 499–518.

Stout, Rowland. 1996. Things That Happen Because They Should. Oxford: Oxford University Press.

Sumpter, David J. T. and Madeleine Beekman. 2003. ‘From Nonlinearity to Optimality: Pheromone Trail Foraging by Ants’. Animal Behaviour 66 (2): 273–80. doi:10.1006/anbe.2003.2224.

Talwar, Victoria, Heidi M. Gordon and Kang Lee. 2007. ‘Lying in the Elementary School Years’. Developmental Psychology 43 (3): 804–10. doi:10.1037/0012-1649.43.3.804.

Tollefsen, Deborah. 2005. ‘Let’s Pretend: Children and Joint Action’. Philosophy of the Social Sciences 35 (75): 74–97.

Tomasello, Michael. 2008. Origins of Human Communication. Cambridge, MA: MIT Press.

Tomasello, Michael and Malinda Carpenter. 2007. ‘Shared Intentionality’. Developmental Science 10 (1): 121–5.

Tomasello, Michael, Malinda Carpenter, Josep Call, Tanya Behne and Henrike Moll. 2005. ‘Understanding and Sharing Intentions: The Origins of Cultural Cognition’. Behavioral and Brain Sciences 28: 675–735.

Träuble, Birgit, Vesna Marinović and Sabina Pauen. 2010. ‘Early Theory of Mind Competencies: Do Infants Understand Others’ Beliefs?’ Infancy 15 (4): 434–44. doi:10.1111/j.1532-7078.2009.00025.x.

Trevarthen, Colwyn. 1980. ‘The Foundations of Intersubjectivity: Development of Interpersonal and Cooperative Understanding in Infants’. In The Social Foundations of Language and Thought: Essays in Honor of Jerome S. Bruner, edited by David Olson, Jeremy M. Anglin and Jerome S. Bruner. New York: Norton.

Triana, Estrella and Robert Pasnak. 1981. ‘Object Permanence in Cats and Dogs’. Animal Learning & Behavior 9 (1): 135–9. doi:10.3758/BF03212035.

Tuomela, Raimo. 2002. ‘Collective Goals and Communicative Action’. Journal of Philosophical Research 28: 29–64.

Tuomela, Raimo. 2005. ‘We-Intentions Revisited’. Philosophical Studies 125 (3): 327–69. doi:10.1007/s11098-005-7781-1.

Tuomela, Raimo and Kaarlo Miller. 1988. ‘We-Intentions’. Philosophical Studies 53 (3): 367–89. doi:10.1007/BF00353512.

Tye, Michael. 1995. Ten Problems of Consciousness. Cambridge, MA: MIT Press.

Uithol, Sebo and Markus Paulus. 2014. ‘What Do Infants Understand of Others’ Action? A Theoretical Account of Early Social Cognition’. Psychological Research 78 (5): 609–22. doi:10.1007/s00426-013-0519-3.

van Buren, Benjamin, Stefan Uddenberg and Brian J. Scholl. 2016. ‘The Automaticity of Perceiving Animacy: Goal-Directed Motion in Simple Shapes Influences Visuomotor Behavior Even When Task-Irrelevant’. Psychonomic Bulletin & Review 23 (3): 797–802. doi:10.3758/s13423-015-0966-5.

van der Wel, Robrecht P. R. D., Natalie Sebanz and Guenther Knoblich. 2014. ‘Do People Automatically Track Others’ Beliefs? Evidence from a Continuous Measure’. Cognition 130 (1): 128–33. doi:10.1016/j.cognition.2013.10.004

Velleman, David. 1997. ‘How to Share an Intention’. Philosophy and Phenomenological Research 57 (1): 29–50.

Velleman, David. 2000. The Possibility of Practical Reason. Oxford: Oxford University Press.

Wan, Lulu, Zoltán Dienes and Xiaolan Fu. 2008. ‘Intentional Control Based on Familiarity in Artificial Grammar Learning’. Consciousness and Cognition 17 (4): 1209–18.

Wang, Bo, Nur Shafiqah Abdul Hadi and Jason Low. 2015. ‘Limits on Efficient Human Mindreading: Convergence Across Chinese Adults and Semai Children’. British Journal of Psychology 106 (4): 724–40. doi:10.1111/bjop.12121.

Wang, Bo, Jason Low, Zhang Jing and Qu Qinghua. 2012. ‘Chinese Preschoolers’ Implicit and Explicit False-Belief Understanding’. British Journal of Developmental Psychology 30 (1): 123–40. doi:10.1111/j.2044-835X.2011.02052.x.

Wang, Su-hua, Renée Baillargeon and Laura Brueckner. 2004. ‘Young Infants’ Reasoning About Hidden Objects: Evidence from Violation-of-Expectation Tasks with Test Trials Only’. Cognition 93 (3): 167–98. doi:10.1016/j.cognition.2003.09.012.

Warneken, Felix, Francis Chen and Michael Tomasello. 2006. ‘Cooperative Activities in Young Children and Chimpanzees’. Child Development 77 (3): 640–63.

Warneken, Felix, Maria Gräfenhain and Michael Tomasello. 2012. ‘Collaborative Partner or Social Tool? New Evidence for Young Children’s Understanding of Joint Intentions in Collaborative Activities’. Developmental Science 15 (1): 54–61. doi:10.1111/j.1467-7687.2011.01107.x.

Warneken, Felix, Jasmin Steinwender, Katharina Hamann and Michael Tomasello. 2014. ‘Young Children’s Planning in a Collaborative Problem-Solving Task’. Cognitive Development 31: 48–58. doi:10.1016/j.cogdev.2014.02.003.

Warneken, Felix and Michael Tomasello. 2006. ‘Altruistic Helping in Human Infants and Young Chimpanzees’. Science 311 (3): 1301–3.

Warneken, Felix and Michael Tomasello. 2007. ‘Helping and Cooperation at 14 Months of Age’. Infancy 11 (3): 271–94. doi:10.1080/15250000701310389.

Wellman, Henry 2018. ‘Theory of Mind: The State of the Art’. European Journal of Developmental Psychology 15 (6): 728–55. doi:10.1080/17405629.2018.1435413

Wellman, Henry and Karen Bartsch. 1994. ‘Before Belief: Children’s Early Psychological Theory’. In Children’s Early Understanding of Mind: Origins and Development, edited by Charlie Lewis and Peter Mitchell. Hove: Erlbaum.

Wellman, Henry and David Cross. 2001. ‘Theory of Mind and Conceptual Change’. Child Development 72 (3): 702–7. doi:10.1111/1467-8624.00309.

Wellman, Henry, David Cross and Julanne Watson. 2001. ‘Meta-Analysis of Theory of Mind Development: The Truth About False-Belief’. Child Development 72 (3): 655–84.

Wellman, Henry, Fuxi Fang and Candida C. Peterson. 2011. ‘Sequential Progressions in a Theory-of-Mind Scale: Longitudinal Perspectives’. Child Development 82 (3): 780–92. doi:10.1111/j.1467-8624.2011.01583.x.

Wellman, Henry and Kristin Lagattuta. 2000. ‘Developing Understandings of Mind’. In Understanding Other Minds: Perspectives from Developmental Cognitive Neuroscience, edited by Simon Baron-Cohen, Helen Tager-Flusberg and Donald J. Cohen. Oxford: Oxford University Press.

Wellman, Henry and David Liu. 2004. ‘Scaling of Theory-of-Mind Tasks’. Child Development 75 (2): 523–41.

Wellman, Henry and Candida C. Peterson. 2013. ‘Deafness, Thought Bubbles, and Theory-of-Mind Development’. Developmental Psychology 49 (12): 2357–67. doi:10.1037/a0032419.

Wellman, Henry, Ann T. Phillips and Thomas Rodriguez. 2000. ‘Young Children’s Understanding of Perception, Desire, and Emotion’. Child Development 71 (4): 895–912.

Wellman, Henry and Jacqueline D. Woolley. 1990. ‘From Simple Desires to Ordinary Beliefs: The Early Development of Everyday Psychology’. Cognition 35 (3): 245–75. doi:10.1016/0010-0277(90)90024-E.

Wenke, Dorit, Stephen M. Fleming and Patrick Haggard. 2010. ‘Subliminal Priming of Actions Influences Sense of Control over Effects of Action’. Cognition 115 (1): 26–38. doi:10.1016/j.cognition.2009.10.016.

Whittlesea, Bruce W. A. 1993. ‘Illusions of Familiarity’. Journal of Experimental Psychology: Learning, Memory, and Cognition 19 (6): 1235–53.

Whittlesea, Bruce W. A. and Lisa D. Williams. 1998. ‘Why Do Strangers Feel Familiar, but Friends Don’t? A Discrepancy-Attribution Account of Feelings of Familiarity’. Acta Psychologica 98 (2–3): 141–65.

Wilcox, Teresa. 1999. ‘Object Individuation: Infants’ Use of Shape, Size, Pattern, and Color’. Cognition 72 (2): 125–66. doi:10.1016/S0010-0277(99)00035-9.

Wilcox, Teresa and Catherine Chapa. 2002. ‘Infants’ Reasoning About Opaque and Transparent Occluders in an Individuation Task’. Cognition 85 (1): B1–B10. doi:10.1016/S0010-0277(02)00055-0.

Wilson, J. A. and J. O. Robinson. 1986. ‘The Impossibly Twisted Pulfrich Pendulum’. Perception 15 (4): 503–4. doi:10.1068/p150503.

Wimmer, Heinz and Michael Hartl. 1991. ‘Against the Cartesian View on Mind: Young Children’s Difficulty with Own False Beliefs’. British Journal of Developmental Psychology 9: 125–38.

Wimmer, Heinz and Heinz Mayringer. 1998. ‘False Belief Understanding in Young Children: Explanations Do Not Develop Before Predictions’. International Journal of Behavioral Development 22 (2): 403–22.

Wimmer, Heinz and Josef Perner. 1983. ‘Beliefs About Beliefs: Representation and Constraining Function of Wrong Beliefs in Young Children’s Understanding of Deception’. Cognition 13: 103–28.

Wolpert, Daniel M., R. Chris Miall and Mitsuo Kawato. 1998. ‘Internal Models in the Cerebellum’. Trends in Cognitive Sciences 2 (9): 338–47. doi:10.1016/S1364-6613(98)01221-2.

Woodward, Amanda L. 1998. ‘Infants Selectively Encode the Goal Object of an Actor’s Reach’. Cognition 69: 1–34.

Woodward, Amanda L. 2009. ‘Infants’ Grasp of Others’ Intentions’. Current Directions in Psychological Science 18 (1): 53–7. doi:10.1111/j.1467-8721.2009.01605.x.

Woodward, Amanda L. and Sarah A. Gerson. 2014. ‘Mirroring and the Development of Action Understanding’. Philosophical Transactions of the Royal Society, B: Biological Sciences 369 (1644): 20130181. doi:10.1098/rstb.2013.0181.

Woodward, Amanda L., Jessica A. Sommerville and Jose J. Guajardo. 2001. ‘Making Sense of Human Behavior: Action Parsing and Intentional Inference’. In Intentions and Intentionality, edited by Bertram F. Malle, Louis Moses and Dare Baldwin, 149–69. Cambridge, MA: MIT Press.

Xu, Fei and Susan Carey. 1996. ‘Infants’ Metaphysics: The Case of Numerical Identity’. Cognitive Psychology 30 (2): 111–53. doi:10.1006/cogp.1996.0005.

Yott, Jessica and Diane Poulin-Dubois. 2016. ‘Are Infants’ Theory of Mind Abilities Well Integrated? Implicit Understanding of Intentions, Desires, and Beliefs’. Journal of Cognition and Development. doi:10.1080/15248372.2015.1086771.

Zhang, Wei and David A. Rosenbaum. 2007. ‘Planning for Manual Positioning: The End-State Comfort Effect for Manual Abduction–Adduction’. Experimental Brain Research 184 (3): 383–9. doi:10.1007/s00221-007-1106-x.

Zwickel, Jan, Sarah J. White, Devorah Coniston, Atsushi Senju and Uta Frith. 2011. ‘Exploring the Building Blocks of Social Cognition: Spontaneous Agency Perception and Visual Perspective Taking in Autism’. Social Cognitive and Affective Neuroscience 6 (5): 564–71. doi:10.1093/scan/nsq088.

Footnotes

  1. This is approximately the story in Plato’s Phaedrus. I might have made up the bit about the traffic accident. ↩︎

  2. I adapt the term ‘inferential integration’ from Stich’s discussion of beliefs. According to him, for beliefs to be inferentially integrated is for there to be ‘generally a huge number of inferential paths via which a given belief can lead to most any other’ (Stich 1978, 506). ↩︎

  3. There is a further respect in which knowledge contrasts with perceptual and motor representations. Knowledge is a pre-theoretical notion which features in social, legal and ethical contexts. By contrast, perceptual and motor representations are theoretical postulates. Their usefulness hinges on their roles in the best available scientific theories of perceiving and acting. In cognitive science, phenomena associated with knowledge are things to be explained whereas perceptual and motor representations are things which explain. ↩︎

  4. Distinguishing perceptual from motor representations will matter in Chapters 7 and 10. ↩︎

  5. This is a simplification; see Sirois and Mareschal (2002) for a more detailed discussion of habituation. ↩︎

  6. Spelke (1990, 50). This principle might need to be refined to accommodate cases where objects are stuck together—indeed, we can regard all of Spelke’s principles as initial guesses subject to revision as more is discovered about cognitive processes concerning physical objects. ↩︎

  7. It is hard to be sure whether Spelke or others endorse the Simple View because there is always room for uncertainty about what they mean by terms like ‘knowledge’. As I use the term, if it is locked to an arbitrarily limited range of actions, or if it used in response to an arbitrarily limited range of events, then it is not knowledge (see Section 1.2). ↩︎

  8. For an opposing view, see Schöner and Thelen (2006); for critical discussion of measures involving looking times generally, see Aslin (2007). ↩︎

  9. Wang, Baillargeon and Brueckner (2004) provide evidence that, in addition to recognizing boundaries and patterns of movement characteristic of objects, 4-month-olds also appear sensitive to conditions under which one object should hide another from their view. This indicates that their abilities to represent objects cannot be fully characterized by the principles we have considered. ↩︎

  10. If you read these studies, you will find that some of the authors talk about Piaget’s stages of object permanence, and about visible and invisible displacements. For our purposes, few of these details matter; the main thing you need to know is just that having object permanence is being able to represent objects as persisting even when they are briefly hidden from your view. ↩︎

  11. According to Davidson, knowledge, intention, desire and all forms of thought depend on belief (1995, 210–11); '[h]aving a belief demands ... appreciating the contrast between true belief and false ' (Davidson 2001, 209); and 'we grasp the concept of truth only when we can communicate the contents—the propositional contents—of the shared experience, and this requires language' (Davidson 1997, 27). ↩︎

  12. Not all nonhumans have difficulties in searching for unperceived objects. Dogs have no difficulty using solidity when searching for an object (Kundey et al. 2010); and young chicks, unlike human infants (Shinskey and Munakata 2001), will search for an object hidden behind a barrier (Chiandetti and Vallortigara 2011). Primates may be special in finding it difficult to search for currently unperceived objects. ↩︎

  13. Since core knowledge is clearly supposed to be distinct from epistemic, perceptual and motor representations, we should really say that core knowledge is supposed to be a fourth type of state. ↩︎

  14. An infelicitous consequence of this definition is that you can have core knowledge of things that are untrue. Perhaps for this reason, Carey (2009, 10) recommends the term 'core cognition' for states of core knowledge. I stick to the term 'core knowledge' because it is so widely used. ↩︎

  15. Kahneman, Treisman and Gibbs (1992, 216), Scholl and Pylyshyn (1999) and Noles, Scholl and Mitroff (2005, 333) all propose a positive answer, but alternative possibilities are sometimes mentioned (for instance, in Odic, Roth and Flombaum 2012, 1078). ↩︎

  16. See Franconeri, Pylyshyn and Scholl (2012). Note that this corrects an earlier argument for a contrary view (Scholl and Pylyshyn 1999). ↩︎

  17. There is some evidence for the view that some abilities to track causal interactions may depend on the errors (or error-like patterns) in assignments of object indexes and the metacognitive feelings these give rise to (see Butterfill 2009, 420ff). Rips (2011, 91ff) offers an objection to this view based on the Pulfrich double pendulum illusion (Wilson and Robinson 1986). This objection could be overcome if assignments of object indexes can diverge from verbal reports of what is seen, or if physical properties such as solidity may sometimes constrain operations on object indexes; and the results of Mitroff, Scholl and Wynn (2005) suggest that both possibilities obtain. As things stand, any link between object indexes and tracking causal interactions must nevertheless be regarded as highly tentative since, as Choi and Scholl (2006, 108) note, the most direct experimental evidence linking abilities to track causal interactions with object indexes comes from a study by Krushke and Fragassi (1996) which, although beautifully designed, is yet to be followed up in print. ↩︎

  18. See Leslie et al. (1998; Scholl and Leslie 1999; Carey and Xu 2001; Scholl 2007). These researchers probably hold differing views on exactly which abilities concerning physical objects can be explained by invoking the CLSTX Conjecture, and are unlikely to endorse my version of the conjecture in all details. ↩︎

  19. See, for example, Johnson et al. (2003). (These experiments involve changes in both distance and temporal duration and were not designed to support a conclusion about how long infants can maintain object indexes.) ↩︎

  20. This assumption is not justified by experimental findings (as far as I know). However, the mundane experience of being able to reach for and grasp things after the lights go off indicates that motor representations of objects survive endarkening. And the fact that endarkening involves destroying the frame of reference indicates that object indexes are unlikely to survive it. ↩︎

  21. This conjecture was suggested by Krisztina Orban (pers. comm., 14 December 2015). ↩︎

  22. Their ironically named paper misses an opportunity to take us 'Beyond the Search Barrier' by interpreting their fascinating findings as showing that 'when task demands are minimal ... young infants search reliably for hidden objects' (McCurry, Wilcox and Woods 2009, 435). Earlier research provides plenty of evidence that task demands are not what prevent 5-month-olds from searching for briefly unperceived objects (see Section 4.1). Further, the experiment of McCurry, Wilcox and Woods (2009) involves no direct manipulation of task demands and so could at most indirectly measure their effects. ↩︎

  23. Compare Dokic (2012, 310): ‘The causal antecedents of noetic feelings can be said to be metacognitive insofar as they involve implicit monitoring mechanisms that are sensitive to non-intentional properties of first-order cognitive processes.’ ↩︎

  24. This illustration is borrowed from Campbell (2002: 133–4); I use it to support a claim weaker than his. ↩︎

  25. For instance, compare Johnston (1992, 222): ‘[j]ustified belief ... is available simply on the basis of visual perception’; Tye (1995, 143–4): ‘Phenomenal character ... stands ready ... to make a direct impact on beliefs’; and Smith (2001, 291): ‘[p]erceptual experiences are ... intrinsically ... belief-inducing’. ↩︎

  26. Those who, like Byrne (2001), hold that phenomenal character is determined by intentional properties would of course reject the existence of sensations in Reid’s sense. They might hold instead that metacognitive feelings are perceptual experiences of the body or of bodily reactions, or that they involve some kind of cognitive intentional object. ↩︎

  27. See Dokic (2012) for an analysis of three approaches to metacognitive feelings (which he calls ‘noetic feelings’). My partial characterization of metacognitive feelings is closest to what he calls the ‘Water Diviner Model’. But I depart from this Model in denying that there is any need to posit intentional contents for metacognitive feelings. This makes my partial characterization of metacognitive feelings closer to Dokic’s view of aesthetic experiences—he argues that aesthetic experiences ‘are non-intentional’ and should be characterized by an adverbial theory (Dokic 2016, 85). I suggest that this is true of metacognitive feelings. ↩︎

  28. An alternative is proposed by Foster and Keane (2015, 79): ‘The MEB theory of surprise posits that: Experienced surprise is a metacognitive assessment of the cognitive work carried out to explain an outcome. Very surprising events are those that are difficult to explain, while less surprising events are those which are easier to explain.’ Foster and Keane investigate reactions to reading about something unexpected, whereas Reisenzein (2000) measures how people experience unexpected events (changes to stimuli while solving a problem). Reisenzein’s study is closer to our interest in infants’ performance on habituation and violation-of-expectation experiments. The truth of either account of surprise, or of an account combining the two insights, would indicate that there is a metacognitive feeling of surprise. ↩︎

  29. An alternative account of how operations involving object indexes influence looking times might start with Smith’s discussion of perceptual anticipations (Smith 2010, 736–9). Butterfill (2015, sec. 4) proposed a way of elaborating on Smith’s notion, relabelling it a ‘phenomenal expectation’. In earlier presentations of this work, I suggested that the objection to Conjecture O could be solved by invoking these phenomenal expectations. I now think that was probably a mistake and that metacognitive feelings provide a better explanation. ↩︎

  30. This is adapted from Pullum and Scholz (2002), who provide a detailed discussion of poverty of stimulus arguments. ↩︎

  31. Compare Spelke (1998, 193): 'If one bases conclusions only on evidence, then I believe that studies of infants suggest that development is not strongly skewed toward either pole of the nativist-empiricist dialogue.' ↩︎

  32. Theorists who assume language users have knowledge of syntax are clearly not referring to the state which is constitutively linked to practical reasoning and to inference, and which is inferentially integrated with other attitudes including belief, desire and intention (see Section 1.2). Chomsky (1965, 8) writes that 'a generative grammar attempts to specify what the speaker actually knows', although the context makes it clear that he cannot be writing about the kind of state identified in Section 1.2. In later work, the linking problem for syntax is dodged by dropping talk of knowledge in favour of an 'I-language' (Chomsky 1995, 14). ↩︎

  33. This is Jackendoff's idea too, although he uses a different term, 'f-knowledge', 'to describe whatever is in speakers' heads that enables them to speak and understand their native language(s)' (2003b, 652; see Jackendoff 2003a, 29ff). ↩︎

  34. Few would agree; see Scholl (2005) for one opposing perspective. ↩︎

  35. There may be very liberal notions of mental representation on which tracking something is sufficient for representing it. Cases like that of the ants demonstrate the value of distinguishing tracking from representing. ↩︎

  36. For an attempt to argue that infants do track intentions, see Woodward (2009, 54–5); in my view. According to Uithol and Paulus (2014, 618), ‘Intention attribution seems to be dependent on language capacities [and] does not emerge until language is sufficiently developed.’ However they support this claim by mentioning just one paper which does not appear to be about intention at all. ↩︎

  37. Or, as they phrase it, ‘an action can be explained by a goal state if, and only if, it is seen as the most justifiable action towards that goal state that is available within the constraints of reality’ (Csibra and Gergely 1998, 255; Csibra 2003). ↩︎

  38. Except possibly where no means of acting successfully are available to the agent, given the current constraints on her possibilities for action. ↩︎

  39. See Paulus et al. (2011, 981): ‘Our results provide evidence that infants do not yet predict actions of others based on the principle of rational action but rather rely on frequency information in forming action predictions.’ ↩︎

  40. At this point, however, it might be objected that the anticipatory looking in question is entirely driven by statistical regularities in the ways objects are used and never by information about the goals of actions (compare Hunnius and Bekkering 2010). Later we will encounter an argument against this objection and a defence of the view that anticipatory looking in 9-month-olds can indeed be driven by information about goals (see Section 11.2). ↩︎

  41. Recent attempts to address this issue (aside from Section 11.4) include Uithol and Paulus (2014) and Gredebäck and Daum (2015). ↩︎

  42. Melzer, Prinz and Daum (2012) argue that, for contralateral grasping, action performance and goal tracking are unrelated at six months of age and only become correlated later in the first year of life. Note also that anticipatory looking to the targets of action is not a pure indicator of goal tracking but is clearly also driven by statistical regularities (as Eshuis, Coventry and Vulchanova 2009; Green et al. 2016) among others argue; see also Section 10.5). ↩︎

  43. This finding is especially puzzling as we have also seen (in Section 10.2) that infants in the first year of life do appear to provide evidence of goal tracking when confronted with scenarios involving self-propelled balls (Csibra and Gergely 1998) or cartoon fish (Daum et al. 2012). ↩︎

  44. Versions of the Motor Theory of Goal Tracking have been defended under various names by many researchers including Falck-Ytter, Gredeback and Hofsten (2006), Kanakogi and Itakura (2011), Ambrosini et al. (2013), Gredebäck and Melinder (2010, 204) and Green et al. (2016). Gredebäck and Falck-Ytter (2015) offer a detailed review of research with adults and infants. Note that researchers differ on how what I am calling the Motor Theory of Goal Tracking relates to the Teleological Stance; some regard the Motor Theory as a competitor to the Teleological Stance rather than (as I suggest in Section 11.3) a further specification of it. ↩︎

  45. Csibra (2008) and Southgate (2013) offer alternative accounts of the relation between the Teleological Stance and the role of motor representations in observing actions. On both their views, motor representations enable predicting joint displacements, bodily configurations and the sensory effects of actions once a goal has been identified (as the Motor Theory also claims); but, contrary to the Motor Theory, they do not matter for identifying goals in the first place. ↩︎

  46. This conjecture is inspired by Gredebäck and Falck-Ytter (2015), Hunnius and Bekkering (2014) and Woodward and Gerson (2014) among others. These authors have interestingly different theoretical positions and would be unlikely to endorse the conjecture, for good reasons. However, they all provide considerations which motivate considering this conjecture. ↩︎

  47. Note that this does not entail rejecting the Motor Theory of Goal Tracking, which merely states that some pure goal tracking involves only motor representations and processes. ↩︎

  48. This line of investigation builds on earlier work by Dittrich and Lea (1994). Those authors use the term ‘goal’ for target. ↩︎

  49. Schlottmann and Ray (2010; Scholl and Tremoulet 2000) all claim that perceptual animacy is a matter of, or involves, tracking goals. ↩︎

  50. This conjecture refines and extends Gredebäck and Melinder's (2010) dual process account, which they offer to explain their findings. There are some differences: they take the Teleological Stance to be an alternative to what I label the Motor Theory of Goal Tracking rather than something presupposed by the Motor Theory; and they take the Teleological Stance to describe what I conjecture are the effects of perceptual animacy (2010, 205). ↩︎

  51. According to an influential definition offered by Premack and Woodruff (1978, 515), for an individual to have a theory of mind it is for her to ‘impute mental states to himself and to others’ (my italics). I have slightly relaxed their definition by changing their ‘and’ to ‘or’ in order to allow for the possibility that there are mindreaders who can identify others’ but not their own mental states. ↩︎

  52. The restriction to typically developing children is necessary because there is evidence that some children either do not make this transition or do so significantly later in childhood. Such children include individuals on the autistic spectrum and deaf children born to hearing parents (Peterson and Siegal 2000). ↩︎

  53. There are, however, studies which find a relation between performance on tasks suitable for infants and tasks used with older children (for example, Meristo, Strid and Hjelmquist 2016). ↩︎

  54. Meristo et al. (2012; Meristo, Strid and Hjelmquist 2016). ↩︎

  55. Rubio-Fernández and Geurts (2012) and Rubio-Fernández (2013) offer a view along these lines. ↩︎

  56. What happens if you attempt to make answering a question about what Maxi, say, believes easier by having Maxi say aloud what he believes before asking the question? When Riggs and Robinson (1995, experiment 2) did this, they found it made no difference. Even though 3- and 4-year-olds only had to repeat something they had just heard to count as passing, and even though they could report what was said when asked about it, they continued to answer questions about belief as if there was no such thing as false belief. ↩︎

  57. See Wellman and Bartsch (1994; Wellman and Lagattuta 2000; Doherty 2011; Roessler and Perner 2013). Doherty, Perner and Roessler all favour the view that 2- or 3-year-olds have an entirely non-mental model of action. However, there is evidence that children understand something of perception, desire, emotion and guesses and their interactions before they can pass standard false belief ↩︎

  58. There is more on using theories to characterize models of minds and actions in Sections 14.5 and 14.6. ↩︎

  59. But which false belief tasks are ‘of the kind 2- and 3-year-olds tend to fail’? This turns out to be an unexpectedly tricky question. See Sections 13.5 and 14.3. ↩︎

  60. Versions of this task have been implemented by Kovács, Téglás and Endress (2010; van der Wel, Sebanz and Knoblich 2014; Edwards and Low 2017). ↩︎

  61. Phillips et al. (2015) offer a detailed challenge to the methods of this study. Although some find the critique convincing (for example, Schneider, Slaughter and Dux 2017, 27), it is striking that follow-up work has successfully found converging results (van der Wel, Sebanz and Knoblich 2014; Edwards and Low 2017), and that a direct attempt to test the challenger account reports findings that support Kovács, Téglás and Endress’s (2010) original interpretation of their findings (El Kaddouri et al. 2019). On balance, Kovács, Téglás and Endress (2010) seems to be passing the replication test with flying colours. ↩︎

  62. I once spent an evening explaining it to Celia Heyes, who told me it was the kind of idea only a philosopher would come up with. I’m not sure this was intended as a compliment. ↩︎

  63. Although they do express their view in terms of belief mirroring, it is possible to interpret Perner and Roessler (2012) as making a suggestion along roughly these lines. ↩︎

  64. Suitable tasks include those which require infants to interpret a pointing gesture (Carpenter, Call and Tomasello 2002; Southgate, Chevallier and Csibra 2010) or a request (Buttelmann, Suhrke and Buttelmann 2015); but we should be cautious at present in using these findings as some attempts to replicate these findings have failed (see Dörrenberg, Rakoczy and Liszkowski 2018 and Kulke and Rakoczy 2018). ↩︎

  65. For an ambitious, relatively detailed attempt involving behavioural patterns, see Ruffman (2014). An alternative approach is suggested by Heyes (2014), who conjectures that infants’ performance on a range of false belief tasks is driven by effects such as retroactive interference and so does not involve any model of minds and actions. ↩︎

  66. Helming, Strickland and Jacob (2015) attempt to explain why children fail A-tasks by appeal to pragmatic considerations. The considerations that some non-A-tasks include communicative prompts and that not all A-tasks need include them are a problem for their attempted explanation. ↩︎

  67. The term ‘declarative expression about belief’ should be understood broadly to include verbal predictions of actions and emotions where the prediction involves ascribing a false belief, and also to include predictions which are expressed by, say, moving a pointer, lifting a flap or pointing to a picture. So making a declarative expression about belief need not involve direct mention of beliefs nor need it involve any verbal or linguistic declaration. ↩︎

  68. Note that this question is not equivalent to asking why there is a gap of months or years between success on non-A-tasks and success on A-tasks. The question arises not simply because A-tasks are in some sense more difficult than other false belief tasks. The question arises because typically developing children’s performance on A-tasks undergoes a striking change over months or years, whereas there appears to be no corresponding age-related change in performance on other false belief tasks. Put statistically, if you compare performance on an A-task with performance on a non-A-task you should observe an interaction of age and task type on performance. ↩︎

  69. A better but more complex analogy for this alternative is Eriksen and Eriksen’s (1974) flanker task. ↩︎

  70. You can find this view in several papers, including any of Leslie and Polizzi (1998; Leslie 2000; Leslie, German and Polizzi 2005). For the application of these ideas to the Mindreading Puzzle, see Scott and Baillargeon (2009), Scott, Baillargeon, Song and Leslie (2010), Baillargeon, Scott and He (2010) or Baillargeon et al. (2015, Section 1-V). I am ignoring some potentially confusing differences between these authors. For instance, Scott, Baillargeon, Song and Leslie (2010, 390–1) and ↩︎

  71. You might wonder whether the selection operation that Leslie et al. focus on is really extraneous to ascribing beliefs. Doherty (1999) pursues this question. ↩︎

  72. These point are often neglected. For example, Baillargeon et al. (2015) assert, incorrectly, that ‘limited executive-function skills … explain young children’s failure at these tasks [i.e. A-tasks]’. ↩︎

  73. One way of refining Leslie et al.’s view might be to drop the claim that inhibition is required when selecting between two contents (the content of the false belief and the content a corresponding true belief would have). Of course, this would then leave open the question of what selection (assuming it is required at all) demands that typical children do not possess until around 4 or 5 years of age. ↩︎

  74. One complication in interpreting these authors’ views is that they appear to talk about different kinds of selection without explicitly distinguishing them or explaining their relation. There is selection between different contents beliefs can have (this is the focus of earlier papers by Leslie et al.) and there is selection between different possible responses to a question or communicative prompt. Here I focus on selection between contents. The idea that children who can track beliefs fail some A-tasks because they have difficulty selecting among possible responses to a question relates to Carruthers’ view; this will be discussed in Section 13.7. ↩︎

  75. Baillargeon et al. (2015, Section 1-V.1) make a similar suggestion. They appear to be endorsing a disjunctive view, on which difficulties of selection and inhibition explains why some children fail some A-tasks (see Section 13.6) whereas something like Carruthers’ view explains performance on other A-tasks. Carruthers (2015a, Section 1.2) describes his own view as ‘consistent with, but somewhat broader than’ that of Baillargeon, Scott and He (2010), whereas I am presenting his view as an alternative to that of Baillargeon, Scott and He (2010). This is because Baillargeon, Scott and He’s and Carruthers’ core ideas are importantly different, even if potentially complementary. (Carruthers also appears not to have considered Baillargeon, Scott and He’s views in detail, for he misdescribes them as concerned with ‘language-involving false belief-tasks’.) ↩︎

  76. There is a temptation to suppose that belief is special because only representing beliefs inconsistent with your own current beliefs involves switching perspectives. But switching perspectives is no less involved in representing other mental states like desires that are incompatible with your own (cf. Rakoczy, Warneken and Tomasello 2007). ↩︎

  77. It would be a mistake to take for granted that A-tasks which rely more heavily on language or communication are somehow more demanding. Given the currently available evidence, the opposite assumption is equally plausible (see Moeller and Schick 2006, 757; Hollebrandse, Hout and Hendriks 2012, 329). ↩︎

  78. Although Schneider is among the authors of one failed attempt to replicate these findings (Kulke et al. 2018), two earlier reports provide successful conceptual replications (Schneider et al. 2012; Schneider et al. 2012). Adding to the puzzle, Naughtin et al. (2017) use essentially the same paradigm as Schneider, Nott and Dux (2014) and found evidence for belief tracking in brain but not behavioural measures. ↩︎

  79. Carruthers (2015b, 9) objects (following Cohen and German 2009) that these experiments are 'not really about encoding belief but recalling it'. But this objection is already answered by Back and Apperly (2010, 56). ↩︎

  80. Or it might turn out that there is no good evidence for the simple dual process theory of mindreading after all. In that case we shall need a different the solution to the mindreading puzzle. ↩︎

  81. See, for example, McCloskey, Washburn and Felch (1983; Kozhevnikov and Hegarty 2001; Oberle et al. 2005). ↩︎

  82. For examples, see Bratman (1987) on intention or Velleman (2000, Chapter 11) on belief. ↩︎

  83. See Braddon-Mitchell and Jackson (1996, 163): 'what is inside our heads should be thought of as more like maps than sentences'. ↩︎

  84. See, for example, Davidson (1963, 1967; Bratman 1987). Philosophers disagree on many features of the canonical theory of the mental, such as whether it treats mental states as intrinsically normative (see, for example, Dretske 2000; Boghossian 2003). ↩︎

  85. Researchers have recently begun to carefully test predictions of this conjecture. Some of the results are surprising, to philosophers at least—influential counterexamples to the claim that knowledge is justified true belief are not regarded as counterexamples at all (for example, Starmans and Friedman 2012; Nagel, Mar and San Juan 2013; Starmans and Friedman 2013). Perhaps philosophers' notions of knowledge differ from those implicit in everyday mindreading. ↩︎

  86. Carruthers (2013, 160) and (2015a, 16) objects that this is no obstacle. His objection rests on the assumption that a thinker who can have a belief whose content we would specify using a particular proposition has thereby represented that proposition and can perform operations using it. But this assumption is not obviously correct (see Matthews 2007, for an overview). ↩︎

  87. The construction of minimal theory of mind described here in barest outline is due to Butterfill and Apperly (2013). It is incomplete in many ways and perhaps inadequate (Christensen and Michael 2016), although it can be improved by integrating it with theories of goal ascription considered in Chapter 10 (as Butterfill and Apperly 2016 explain). ↩︎

  88. See, for example, Heal (2002; McGeer 1996). ↩︎

  89. Encountering and registration were defined in Section 14.6. ↩︎

  90. It is important to distinguish between testing a hypothesis (first make the prediction, then do the experiment) and fitting a hypothesis to existing results (first study the findings, then formulate a hypothesis). There will always be alternative post hoc explanations for evidence that confirms a prediction. ↩︎

  91. Note that speed-accuracy trade-offs carry little argumentative weight. They play no essential role in characterizing mindreading processes, nor in justifying the acceptance of hypothesis. They serve merely to motivate testing hypotheses. ↩︎

  92. For a particular response, let A be the probability that automatic mindreading influences this response and C the probability that non-automatic mindreading influences it. To say that automatic mindreading processes dominate on a particular task is to say that the task involves measuring a response for which the ratio of A(1–C) to C is large enough to enable us to measure the effects of automatic mindreading. ↩︎

  93. Some researchers have used the term 'identity' not for numerical identity but in referring to the category of an object (for example, whether it is a fish or a sponge: Buttlemann and Kovács, n.d.; Buttlemann, Suhrke and Buttlemann 2015). As the signature limit under discussion concerns numerical identity, not the category or function of an object, these studies are not directly relevant here (although the findings are fascinating in their own right). ↩︎

  94. For example, Schneider, Nott and Dux's (2014, 46) findings indicate that anticipatory looking during one interval reflects automatic mindreading, whereas anticipatory looking during another, later interval reflects non-automatic mindreading. ↩︎

  95. Low (2010, 613) appears to endorse this conclusion. ↩︎

  96. According to Pyers and Senghas (2009, 810), 'Language is indeed a necessary prerequisite, one that cannot be replaced by even 25 years of social experience. Adults who had no congenital cognitive deficits, but whose language was incomplete, failed to fully understand the beliefs of others.' See also Meristo et al. (2007, 1166):

    The expression of [mindreading] in native-signing deaf children may also depend on children's continuing exposure to opportunities for monitoring the nature of conversational input about mental states and its implications for evaluating beliefs and other mental states as true or false. ↩︎

  97. Intriguingly, there is also evidence that infants born deaf into hearing families do not pass non-A-tasks (Meristo et al. 2012; see also Morgan et al. 2014). ↩︎

  98. In a later study with older children, Warneken, Gräfenhain and Tomasello (2012) showed that children will re-engage a partner even when it would be possible for them to achieve the goal without any contribution from the partner. ↩︎

  99. In the passage just cited, Brownell identifies the need for an operational characterization of joint action. What follows is merely an attempt to provide a theoretical characterization, albeit one simple enough to support the further work involved in providing an operational characterization. ↩︎

  100. The collective and distributive readings are not the only possibilities, but they are the only ones that need concern us here. ↩︎

  101. There are some tricky cases in which different performers have expectations concerning slightly different actions. Since we are considering a merely sufficient condition for joint action, we are not required to resolve these. ↩︎

  102. For instance, Bratman (2014, 57) defends including a requirement on common knowledge by asserting that ‘public access to the shared intention will normally be involved in further thought that is characteristic of shared intention, as when we plan together how to carry out our shared intention’. This assertion appears to justify the claim that common knowledge accompanies shared intention rather than the claim that common knowledge is an essential feature of shared intention. ↩︎

  103. According to Gopnik and Meltzoff (1997), humans’ development is a consequence of their pursuing investigations in roughly the way scientists do (see Gopnik 1996, for a shorter statement of the view for philosophers). I suspect infants of doing a bit of ethics and legal work on the side. ↩︎

  104. The exception is mindreading: we did not yet see evidence linking mindreading to perceptual or motoric processes. My guess is that this is because the research is yet to be done. We may eventually discover that the earliest developing abilities to track mental states are in fact a consequence of a combination of broadly perceptual and broadly motoric processes. ↩︎

  105. You might suspect Davidson’s target is not scientists at all, but the context indicates otherwise. Just after this passage (p. 128), he reports, ‘I am thankful that I am not in the field of developmental psychology.’ ↩︎

  106. But if you invite her to dinner, be careful who you put next to her. ↩︎

Command Palette
Search for a command to run