Resources#

dacy.resources.names#

Helper functions for loading name dictionaries for person augmentation.

dacy.resources.names.danish_names()[source]#

Returns a dictionary of Danish names.

Returns

A dictionary of Danish names containing the keys “first_name” and “last_name”. The list is derived from Danmarks statistik (2021).

Return type

dict[str, list[str]]

Example

>>> from dacy.resources import danish_names
>>> names = danish_names()
>>> names["first_name"]
>>> names["last_name"]
dacy.resources.names.female_names()[source]#

Returns a dictionary of Danish female names.

Returns

A dictionary of names containing the keys “first_name” and “last_name”. The list is derived from Danmarks statistik (2021).

Return type

dict[str, list[str]]

Example

>>> from dacy.resources import female_names
>>> names = female_names()
>>> names["first_name"]
>>> names["last_name"]
dacy.resources.names.load_names(min_count=0, ethnicity=None, gender=None, min_prop_gender=0)[source]#

Loads the names lookup table. Danish are from Danmarks statistik (2021). Muslim names are from Meldgaard (2005), https://nors.ku.dk/publikationer/webpublikationer/muslimske_fornavne/.

Parameters
  • min_count (int) – Minimum number of occurrences of the name for it to be included. Defaults to 0.

  • ethnicity (str | None) – Which ethnicity should be included. None indicate all is included. Options include “muslim”, “danish”. Defaults to None.

  • gender (str | None) – Which gender should be included. None indicate all is included. Options include “male”, “female”. Defaults to None.

  • min_prop_gender (float) – Minimum probability of a name being a given gender. The probability of a given name being a specific gender is based on the proportion of people with the given name of that gender. Only used when gender is set. Defaults to 0.

Returns

A dictionary of names containing the keys “first_name” and “last_name”.

Return type

dict[str, list[str]]

dacy.resources.names.male_names()[source]#

Returns a dictionary of Danish male names.

Returns

A dictionary of names containing the keys “first_name” and “last_name”. The list is derived from Danmarks statistik (2021).

Return type

dict[str, list[str]]

Example

>>> from dacy.resources import male_names
>>> names = male_names()
>>> names["first_name"]
>>> names["last_name"]
dacy.resources.names.muslim_names()[source]#

Returns a dictionary of Muslim names.

Returns

A dictionary of Muslim names containing the keys “first_name” and “last_name”. The list is derived from Meldgaard (2005), https://nors.ku.dk/publikationer/webpublikationer/muslimske_fornavne/.

Return type

dict[str, list[str]]

Example

>>> from dacy.resources import muslim_names
>>> names = muslim_names()
>>> names["first_name"]
>>> names["last_name"]

dacy.resources.dictionaries#

Helper functions for loading Danish dictionaries.

dacy.resources.dictionaries.load_ods_fullforms(redownload=False)[source]#

Loads the full-form list from ODS, a historical Danish dictionary.

Ordbog over det danske Sprog (ODS) covers Danish from approximately 1700 to 1950. The August 2020 full-form list contains about 1.3 million entries, including inflected forms. Original spelling and capitalization are preserved.

DSL’s terms of use allow reuse, modification and redistribution, including commercial use, but prohibit publishing a dictionary or a product competing with DSL’s products. Attribution to DSL is requested. Downloading accepts these terms.

Parameters

redownload (bool) – Download again even if the file is cached. Defaults to False.

Returns

A DataFrame with five string columns.

  • form: The word form, including inflected forms.

  • headword: The dictionary headword associated with the form.

  • homograph: The number distinguishing dictionary entries with the same headword spelling; an empty string when absent.

  • pos: The part of speech, using ODS abbreviations.

  • id: The identifier of the dictionary entry in ODS.

Multiple entries for the same form are retained.

Return type

DataFrame

Example

Look up the inflected form “Hesten” (“the horse”) and its headword “Hest” (“horse”):

>>> from dacy.resources import load_ods_fullforms
>>> entries = load_ods_fullforms()
>>> print(entries.loc[entries["form"] == "Hesten"].to_string(index=False))
  form headword homograph pos       id
Hesten     Hest           sb. 60137181
>>> wordforms = set(entries["form"])