.. _plurals:
Plural Normalisation
======================
This is the whole point of this fork of the code. Normalizing plurals is hard
and there are lots of weird edge cases.
New Behaviour
-------------
The new version uses the Python `inflect `_
library which can do a lot more than find plurals. In our case, we just use it
to take a word that *might* be plural, and either make it singular, or leave it
alone.
If ``normalize_plurals`` is ``True``, then the following rules apply:
All instances of a plural are converted to the singular
'''''''''''''''''''''''''''''''''''''''''''''''''''''''
If both the singular and plural version of a word appear in the corpus, all
instances of the plural are converted to the singular for purposes of
weight/frequency. So ``['cat', 'cats', 'cat']`` effectively becomes ``['cat',
'cat', 'cat']``.
If only the plural appears in the corpus, it is unmodified
''''''''''''''''''''''''''''''''''''''''''''''''''''''''''
So ``['cats', 'cats', 'cats']`` is still treated as ``['cats', 'cats',
'cats']``. Singular forms that never appeared in the original will not be
introduced by normalizing. For example ``['mice', 'mice', 'mice']`` will not
become ``['mouse', 'mouse', 'mouse']``, because ``mouse`` never appeared in the
original source text. However, ``['mouse', 'mice', 'mice']`` will become the
equivalent of ``['mouse', 'mouse', 'mouse']``.
Phrases (bigrams) that use the plural word keep it, rather than get made singular
'''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''
So ``['cute cats', 'cat', 'cats']`` will be treated like ``['cute cats', 'cat',
'cat']``. In this example **cats**, when alone, gets normalised to **cat**. But
the phrase **cute cats** will still appear with the plural **cats**.
Known Behaviours
````````````````
**Verbs as Nouns**: The ``inflect`` library only has a function
``singular_noun()`` which will treat every word as if it was a noun. That means
that a word like **is**, when treated like a noun, gets shortened to **i** (as
if we are talking about several letters "i": *'There are 2 Is in nutrition.'*).
Lots of verbs can also be nouns. It means that a phrase like **He walks every
day. He enjoys his walk.** will end up normalizing **walks** to **walk**, even
though **walks** was used as a verb.
Old Behaviour
-------------
The old behavior was described `in the comments `_ of ``tokenization.py``:
Each word is represented by the most common case.
If a word appears with an "s" on the end and without an "s" on the end,
the version with "s" is assumed to be a plural and merged with the
version without "s" (except if the word ends with "ss").
This has the obvious deficiencies of missing a lot of irregular plurals (mouse/mice), plus mishandling some words that look plural but aren't (premises). It works surprisingly well, but not **that** well.