Download Jane: User's Manual

Transcript
Chapter 4
Rule extraction
In this chapter, we are going to look a little bit more closer at the extraction.
In Section 4.1, we roughly explain a typical extraction workflow In Section 4.2, we
look at the options of the training script. For more details of the various options of the
single tools, we present some of the most important ones in Section 4.3 and Section 4.4.
In general, the descriptions are mainly taken from the man pages of the tools. Be
aware that the specific options might change (and there is a non-zero possibility that we
forgot to update this manual), so if in doubt always believe the man/help pages.
4.1
Extraction workflow
The idea of the rule training is that we extract the counts of each rule that we are
interested in, and afterwards normalize them, i.e. compute their relative frequencies.
We typically filter the rules to only those that are needed for the translation. Otherwise,
even for medium-sized corpora the files are getting too large.
Mandatory files are the corpus, consisting of the source and the target training file,
and their alignment. Highly recommended—especially for hierarchical rule extraction—
is the source filter file. We actually extract twice:
In the first run, we generate the actual rule counts. They are filtered with the source
filter file (by using suffix arrays). Now that we know which target counts we will need
for normalization, we run the extraction a second time, filtering with the source filter
file and all rule targets that were generated (by using prefix trees).
With some lexical counts that are typically extracted from the corpus, we are now
able to start the normalization and produce our actual rule table. See Figure 4.1 for a
graphical representation.
4.2
Usage of the training script
Jane provides a shell script which basically performs all the necessary operations as
mentioned above.
43