Download Jane: User's Manual
Transcript
Chapter 4 Rule extraction In this chapter, we are going to look a little bit more closer at the extraction. In Section 4.1, we roughly explain a typical extraction workflow In Section 4.2, we look at the options of the training script. For more details of the various options of the single tools, we present some of the most important ones in Section 4.3 and Section 4.4. In general, the descriptions are mainly taken from the man pages of the tools. Be aware that the specific options might change (and there is a non-zero possibility that we forgot to update this manual), so if in doubt always believe the man/help pages. 4.1 Extraction workflow The idea of the rule training is that we extract the counts of each rule that we are interested in, and afterwards normalize them, i.e. compute their relative frequencies. We typically filter the rules to only those that are needed for the translation. Otherwise, even for medium-sized corpora the files are getting too large. Mandatory files are the corpus, consisting of the source and the target training file, and their alignment. Highly recommended—especially for hierarchical rule extraction— is the source filter file. We actually extract twice: In the first run, we generate the actual rule counts. They are filtered with the source filter file (by using suffix arrays). Now that we know which target counts we will need for normalization, we run the extraction a second time, filtering with the source filter file and all rule targets that were generated (by using prefix trees). With some lexical counts that are typically extracted from the corpus, we are now able to start the normalization and produce our actual rule table. See Figure 4.1 for a graphical representation. 4.2 Usage of the training script Jane provides a shell script which basically performs all the necessary operations as mentioned above. 43