Loading...
Please wait, while we are loading the content...
Similar Documents
Statistical significance tests for machine translation evaluation.
| Content Provider | CiteSeerX |
|---|---|
| Abstract | If two translation systems differ differ in perfor-mance on a test set, can we trust that this indicates a difference in true system quality? To answer this question, we describe bootstrap resampling meth-ods to compute statistical significance of test results, and validate them on the concrete example of the BLEU score. Even for small test sizes of only 300 sentences, our methods may give us assurances that test result differences are real. 1 |
| File Format | |
| Access Restriction | Open |
| Subject Keyword | Statistical Significance Test Machine Translation Evaluation Bleu Score True System Quality Test Set Test Result Difference Concrete Example Statistical Significance Translation System Test Result Small Test Size |
| Content Type | Text |