Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The main reason I posted was in the hope of triggering discussion over the method used to analyse the text.

I've seen users getting a lower score simply by separating out a block of text into numbered paragraphs which would seem to point to quite a simplistic method.

http://ipdraughts.wordpress.com/2012/08/25/cutting-down-on-t...



I'd like to use it as a plugin for email, Reddit Enhancement Suite, and other such things of that nature .. ;)


I'll withstand my statement: model based on a corpus of PR, scholar, licenses and the like texts. If they are into real statistical NLP.

Or just esthetic rules + word dictionary.


If I were to make the software, the corpus of PR, licenses, etc. would be the way I go. But "they did it statistically" doesn't answer the question "what is the model?" There are many different statistical models one could use. My other post has a few things we've figured out.

But I'm starting to think a rule-based lexicon isn't out of the question, given these >1 scores on some texts.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: