.. _plurals: Plural Normalisation ====================== This is the whole point of this fork of the code. Normalizing plurals is hard and there are lots of weird edge cases. New Behaviour ------------- The new version uses the Python `inflect `_ library which can do a lot more than find plurals. In our case, we just use it to take a word that *might* be plural, and either make it singular, or leave it alone. If ``normalize_plurals`` is ``True``, then the following rules apply: All instances of a plural are converted to the singular ''''''''''''''''''''''''''''''''''''''''''''''''''''''' If both the singular and plural version of a word appear in the corpus, all instances of the plural are converted to the singular for purposes of weight/frequency. So ``['cat', 'cats', 'cat']`` effectively becomes ``['cat', 'cat', 'cat']``. If only the plural appears in the corpus, it is unmodified '''''''''''''''''''''''''''''''''''''''''''''''''''''''''' So ``['cats', 'cats', 'cats']`` is still treated as ``['cats', 'cats', 'cats']``. Singular forms that never appeared in the original will not be introduced by normalizing. For example ``['mice', 'mice', 'mice']`` will not become ``['mouse', 'mouse', 'mouse']``, because ``mouse`` never appeared in the original source text. However, ``['mouse', 'mice', 'mice']`` will become the equivalent of ``['mouse', 'mouse', 'mouse']``. Phrases (bigrams) that use the plural word keep it, rather than get made singular ''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''' So ``['cute cats', 'cat', 'cats']`` will be treated like ``['cute cats', 'cat', 'cat']``. In this example **cats**, when alone, gets normalised to **cat**. But the phrase **cute cats** will still appear with the plural **cats**. Known Behaviours ```````````````` **Verbs as Nouns**: The ``inflect`` library only has a function ``singular_noun()`` which will treat every word as if it was a noun. That means that a word like **is**, when treated like a noun, gets shortened to **i** (as if we are talking about several letters "i": *'There are 2 Is in nutrition.'*). Lots of verbs can also be nouns. It means that a phrase like **He walks every day. He enjoys his walk.** will end up normalizing **walks** to **walk**, even though **walks** was used as a verb. Old Behaviour ------------- The old behavior was described `in the comments `_ of ``tokenization.py``: Each word is represented by the most common case. If a word appears with an "s" on the end and without an "s" on the end, the version with "s" is assumed to be a plural and merged with the version without "s" (except if the word ends with "ss"). This has the obvious deficiencies of missing a lot of irregular plurals (mouse/mice), plus mishandling some words that look plural but aren't (premises). It works surprisingly well, but not **that** well.