Monday, September 22, 2008

Colorless green ideas

Here is Chomsky's footnote 4 from Syntactic Structures, two pages after his introduction of the famous Colorless green/Furiously sleep pair, which, he argues, shows that "in any statistical model for grammaticalness, these sentences will be ruled out on identical grounds as equally 'remote' from English."


One might seek to develop a more elaborate relation between statistical and syntactic structure than the simple order of approximation model we have rejected. I would certainly not care to argue that any such relation is unthinkable, but I know of no suggestion to this effect that does not have obvious flaws. Notice, in particular, that for any n, we can find a string whose first n words may occur as the beginning of a grammatical sentence S1 and whose last n words may occur as the ending of some grammatical sentence S2, but where S1 must be distinct from S2. For example, consider the sequences of the form "the man who ... are here," where ... may be a verb phrase of arbitrary length. Notice also that we can have new but perfectly grammatical sequences of word classes, e.g., a sequence of adjectives longer than any ever before produced in the context "I saw a -- house." Various attempts to explain the grammatical-ungrammatical distinction, as in the case of (1), (2), on the basis of frequency of sentence type, order of approximation of word class sequences, etc., will run afoul of numerous facts like these.


Even without rooting through LSLT, we see that Chomsky has immediately qualified his any statistical model, so I think any schism centring around the Colorless... argument is a result of miscommunication more than anything. This is just an argument that you can't get a blood out of a stone, and distributional information per se is to a theory of syntax a mere bloodless stone (contrary to what some people at the time apparently were proposing).

The details of his argument are not totally clear to me, but the argument about agreement seems to be an argument from nested dependencies, thus, against using a finite-state model; the argument from iterated modifiers seems to be an argument against a model in which we retrieve one of a finite set of strings of category labels, and replace each of the labels with a word. And if I've reconstructed that part of it correctly, the full argument perhaps was meant to be basically what Chris said today: if there's something fundamentally wrong with the model of language with respect to which we keep track of distributional information, then although you may get a reasonable approximation, ultimately, more data, whether you keep track of distributions or not, is not going to help. Having distributional information doesn't replace a theory of language; it presupposes one. Implicitly, one might imagine rounding the argument off by saying, Once we've come up with a proper theory of language, then maybe we can find a role for keeping track of distributions. Or maybe we won't need it.

So how does this relate to the contents of the course? Well, although it seems clear that EM is deeply important, and it's clear that finite state mappings of the kind we've been talking about are part of the language faculty, it's not clear to me that the interesting learning problem in, say, phonology, arises when we know the structure of the model and the only parameters are transition/emission probabilities; the problem is learning the structure of the machine. So these examples are admittedly a bit hard to sink my teeth into. Clustering and constructing features, however--the what's important here problems--we can easily cook up problems that look like that in any subfield of linguistics (for example, representations in both phonetics and phonology).

No comments: