about
The IIT Bombay English-Hindi Parallel Corpus [pdf] (arxiv.org)
103 points by aq3cn on Oct 10, 2017 | hide | past | pdf | 13 comments on HN

In plain words: A free collection of matching English and Hindi sentences, gathered from earlier public sources plus newly collected text and cleaned for training translation systems. It holds 1.49 million sentence pairs, nearly half of them newly released, making it the biggest public English-Hindi set.

Abstract · The IIT Bombay English-Hindi Parallel Corpus

We present the IIT Bombay English-Hindi Parallel Corpus. The corpus is a compilation of parallel corpora previously available in the public domain as well as new parallel corpora we collected. The corpus contains 1.49 million parallel segments, of which 694k segments were not previously available in the public domain. The corpus has been pre-processed for machine translation, and we report baseline phrase-based SMT and NMT translation results on this corpus. This corpus has been used in two editions of shared tasks at the Workshop on Asian Language Translation (2016 and 2017). The corpus is freely available for non-commercial research. To the best of our knowledge, this is the largest publicly available English-Hindi parallel corpus.

Anoop Kunchukuttan, Pratik Mehta, Pushpak Bhattacharyya
arXiv:1710.02855 · cs.CL · submitted Oct 8, 2017 · updated May 19, 2018
abstract · pdf · html · accepted for LREC 2018, 4 pages, parallel corpus for English-Hindi machine translation

add comment on HN

Since so long, I have been waiting for Indian universities especially IITs to invest and publish in building such corpora. Being a founder of AI/ML startup, I am surprised at the appalling lack of datasets available to work on Indian problems. Contrast this with Chinese universities where they have built some world class datasets to build NLP solutions in Mandarin. Our sentiment analysis works in 8 different languages but none of it is in Indian languages despite we being in India!
The data set is released as CC-BY-NC and, thus, cannot be used in commercial applications.
And also cannot be used in open-source projects.

CC-BY-NC amounts to saying: you can play around with this in demos and academic projects that no lawyer would ever go after anyway, but you can't use it for real.

Is that entirely a bad thing? If a commercial, for profit wants to enter this field then they can pay for it or licence it?
And what if an open source project wants to enter this field?
For open source commercial, it's the same as for-profit. They'll probably get a discount.

For open source non-profit, it'd still be NC and therefore legal?

"Non-profit" is a whole different kettle of fish than "non-commercial". It doesn't mean you're not selling anything. It doesn't mean you're morally good. It just means you registered for a particular business status with particular restrictions. And it's not what CC is talking about.

A key part of the Open Source Definition is that you do not lose your permission to use and copy the code based on what you do with it. A project that you aren't allowed to use anymore if you start making money from it is not Open Source.

"Non-commercial" code restrictions are more like the thing where you're allowed to look at the Windows source code, if you're an academic and you ask nicely and you won't ever do anything with it.

A lot of universities are open to give a separate commercial license when contacted. They charge for their efforts, which is fair. Whether public universities should charge given we already pay them from our taxes is a different issue. source: I am also cofounder at an AI startup, we often buy licenses to use academic datasets for commercial usage.
I guess it’s not the most demanding requirement at the moment. But happy to see progress.

Btw, where do you work? Can we see your work?

Being it's IIT Bombay I hope for Marathi some day soon.

I know different people have different priorities :-)

Does it matter that it is in Bombay? I thought the prod/student body would be diverse due to to how the intake works.

Nothing against Marathi, I feel like Telugu, Tamil etc would also have equal chances of happening there.

Oh not really, I do agree that languages like Tamil, Telugu, Bengali and others are sort of “more deserving” in that they have more speakers than Marathi. I just don’t speak them and probably never will :-( .

The “Bombayness” was only that Mumbai is in Maharashtra. There really are great folks from all over at that institution. My comment was NOT intended to make any implication about any community over another! I see someone downvoted my comment and might have understandably thought that that might have been my intention.

Actually I was saying the total opposite of what you inferred. I thought that despite being in Mumbai, IIT-B would have a very diverse spread of folks due to the methods of intake(JEE/ all India recruitment of professors). That means Marathis might not be the language of choice, because it's highly likely that the researcher is Tamil/Telugu. Unlike someplace like, say, VJTI, that's likely to have a large Marathi population.

(I didn't downvote btw.)