Marcus, MitchellSantorini, BeatriceMarcinkiewicz, Mary Ann2023-05-222023-05-221993-10-012007-07-11https://repository.upenn.edu/handle/20.500.14332/7145In this paper, we review our experience with constructing one such large annotated corpus--the Penn Treebank, a corpus consisting of over 4.5 million words of American English. During the first three-year phase of the Penn Treebank Project (1989-1992), this corpus has been annotated for part-of-speech (POS) information. In addition, over half of it has been annotated for skeletal syntactic structure.Building a Large Annotated Corpus of English: The Penn TreebankReport