Jump to content

Yezur Wiki:Dictionary structure

From Yezur Wiki

How the Dictionary namespace should be structured, modelled on Earth's Wiktionary (specifically the English Wiktionary, whose "Entry layout" policy is the most developed). This page is the out-of-frame blueprint; the Dictionary entries themselves are in-character. The skeleton below is in use — it documents the form entries take, so that words are added in a structured, precise way.

How Wiktionary structures an entry

One page per written form. The page title is the exact spelling, case-sensitive (polish and Polish are different pages). Every language that uses that spelling shares the single page; a {{also|…}} hatnote at the very top cross-links near-identical spellings (case or diacritic variants).

One level-2 heading per language (==English==, ==Finnish==), in alphabetical order. Everything about the word in that language nests inside its section; languages never mix below level 2. A word's entry can be a two-line stub or a monograph, but it always sits in the same skeleton.

A fixed heading order inside each language section — only the headings that apply, but always in this order and nesting:

==Language==
===Alternative forms===
===Etymology===
===Pronunciation===        (IPA, audio, rhymes, hyphenation)
===Noun===                 (or Verb, Adjective, … — the part-of-speech heading)
Headword line
# Meaning 1
#* Quotations
====Usage notes====
====Synonyms==== / ====Antonyms==== / …
====Derived terms====
====Related terms====
====Descendants====
====Translations====       (English entries only)
===References===
===Anagrams===

Where one spelling covers several unrelated origins, ===Etymology 1===, ===Etymology 2===, … each carry their own part-of-speech sections, pushed one level deeper (====Noun====). Nesting is the key principle: a heading owns everything after it until a heading of the same level.

The headword line. Directly under each part-of-speech heading, a headword template renders the word in bold with its key grammar — language-specific where one exists ({{en-noun}}, which knows English plurals), the generic {{head|<code>|<pos>}} otherwise. It also auto-categorizes the page (Category:English nouns) — editors almost never hand-categorize.

Definitions are numbered wikitext list items (#), one sense per line. Context labels — {{lb|en|nautical|slang}} — prefix a sense and categorize it too. Example sentences ({{ux|…}}) and dated quotations (#*) nest under the sense they illustrate.

Everything is language-coded. The first argument of nearly every template is a language code (en, ang, gem-pro), resolved by one central registry module (their Module:languages) that knows each language's canonical name, script, and family. That single registry is what lets templates link, label, transliterate, and categorize automatically — it is the engine of the whole system.

Linking vs. mentioning. {{l|de|Hund}} links a term in a list (synonyms, derived terms); {{m|de|Hund}} mentions it in running prose (italicized, as in etymologies). Both target the right language's section of the right page.

Etymology templates display text and categorize in one stroke:

  • {{inh|en|ang|hund}} — inherited from an ancestor stage
  • {{bor|en|fr|hôtel}} — borrowed
  • {{der|…}} — derived, when the path is less direct
  • {{cog|de|Hund}} — a cognate in a sister language (display only)
  • {{doublet|en|canine}} — same-source pairs within the language

Non-English entries are glossed, not defined: a Finnish word gets a one-line English translation as its definition, and no Translations section. Translation tables ({{trans-top|gloss}} … {{t|de|Hund|m}} … {{trans-bottom}}, one collapsible table per sense) live only on English entries. English is the hub language — every pair of languages meets through it, which keeps the many-to-many translation problem linear.

Reconstructed terms are not mainspace. Proto-language words live in a dedicated Reconstruction: namespace (Reconstruction:Proto-Germanic/hundaz) and are always cited with a leading asterisk (*hundaz), the historical-linguistics convention for unattested forms.

Inflected forms get thin satellite pages ("hounds — plural of hound"), usually bot-created; the lemma page carries all real content.

What the Yezur Dictionary adopts

The same skeleton, sized for a dictionary starting from zero:

  • Adopt as-is: one page per written form; one L2 section per Yezuri language; the fixed heading order and nesting (including Etymology 1 / Etymology 2 for homographs, from the first entry); numbered senses; glosses in the wiki's prose language — in-world, Traditional English (see below).
  • One Related terms heading, with labelled bullets (operator ruling), in place of Wiktionary's heading per relation (====Synonyms====, ====Antonyms====, ====Derived terms====, ====Related terms====). Every lexical relation gathers under a single ====Related terms==== heading, each bullet opening with the relation it states: * Antonym: {{l|YAP-XX|word}} (“gloss”) — likewise Near-synonym:, Coordinate term:, Derived term: and Compare:. ====Descendants==== keeps its own heading, its bullets labelled by language.
  • Within a definition line, the comma joins and the semicolon explains (operator ruling, 2026-08-04). Items that could each stand alone as a translation of the headword are joined by commas (# fear, be afraid, dread); a semicolon marks the shift from translating to describing, introducing an explanation or elaboration (# nose; the smelling part). The two combine naturally: # seed, grain; the sown thing.
  • Registry first — reusing the wiki's existing code system. Module:Languages is the registry: it maps the same codes the language infoboxes already carry (family-prefixed: YAQ-GA Gaillean, YAQ-HE Hertic; YBS-LO Lonish, YBS-NI Nichana, YBS-PJ Pjany, YBS-UU Uu; plus YXS-EN / YXS-NE for the two forms of English) to display name and Dictionary section heading. Every dictionary template takes the code as an argument, exactly as on Wiktionary. Keep the module in step with the code field of Template:Infobox language.
  • {{d|word|code}} — built (Template:D): links a word to its Dictionary entry, anchored to the language's section through the registry — {{d|dog|YXS-EN}} leads to the English section of Dictionary:dog. Use it wherever prose refers to a word.
  • A minimal starter kit of further general-purpose templates (per-language headword templates can come later, once grammar exists to encode):
    • {{head|code|pos}} — headword line + auto-categorization + a Lexemes row (built, Template:Head)
    • {{l|code|word}} and {{m|code|word}} — link and mention (built)
    • {{lb|code|label…}} — inline sense labels, italic and parenthesised before a sense (built, Template:Lb)
    • {{auphen|word|code}} — automatic pronunciation from the language's ruleset (built, Module:Auphen); plain {{IPA}} for hand-written IPA
    • {{inh}} / {{bor}} / {{cog}} — the etymology trio (still to build)
    • {{ux|code|example}} — usage examples (still to build)
  • Cargo from the first word. Where Wiktionary only categorizes, our {{head}} also files each entry into the Lexemes Cargo table (word, language, part of speech, gloss, …). That is the interlingual payoff the Dictionary exists for: cognate sets, borrowing chains, "every Hertic noun", "all words inherited from Proto-Herta-Olaic" — one query each, self-maintaining. Since the encyclopedia retired its Cargo tables (2026-07), Lexemes is the wiki's only Cargo table — a deliberate exception, kept because the lexicon is the one place where scale outruns hand curation.
  • Reconstructions in the same namespace (operator's call — no Reconstruction: analogue): a reconstructed form is an ordinary Dictionary page, its title the bare form without the asterisk; the proto-language is simply another language section. The headword line carries the leading * marking the form as unattested, as do mentions in running text — but never the page title.
  • Translation boxes — built (Template:Translations, Module:Translations): one box per sense, directly under that sense's definitions in an English lemma section, with a required gloss= naming the sense so that homonyms stay apart. Languages are keyed by registry code and grouped by family through Module:Languages, which also links each word to its entry and stars the reconstructed ones; unrecognized codes show as errors rather than being dropped. Display-only — the box stores no data and is hand-curated, like the family Members tables. English is the hub language, as on Wiktionary.

The English frame

(Recorded from the operator, 2026-07-17; the in-world articles — Engil, the English language on Yezur — are still to be written.)

English is a language spoken on Yezur: a relatively small isolate spoken in the middle of the ocean, on the island nation of Engil. Its origins are nebulous and not well researched — which obscures the large amount of influence it has been under. It has two forms:

  • Traditional English (YXS-EN) — identical to Terran English in all ways except Yezuri semantic coinages. The wiki's prose is Traditional English, as most people on Engil speak it.
  • New English (YXS-NE) — the globalized form: the same language with roughly 3% of the vocabulary swapped for Anglified Yezuri loans or words of different etymology. In Terran terms, the replaced words fall into three main categories:
    1. words English loaned from non-European languages after the cutoff of roughly the 1500s (coffee, zebra, apartheidbeef is fully naturalized and stays; bouillon, c. 1600s, is a little too late and goes);
    2. words the world typically loans from English (computer, internet, hype);
    3. words with distinctly Terran origins (sandwich, rugby, saxophone, pasteurize).

The two forms are not treated as separate languages — comparable to how Wiktionary treats British and American English. One ==English== section serves both (both codes anchor there via the registry); a New English word such as tuku gets its own entry page, labelled as New English within its English section, and cross-referenced from its Traditional counterpart zebra — and vice versa. In article prose: "Zebras (New English: tuku) are equine animals…".

Settled decisions

  • Native-script page titles, with romanized forms as redirects. Yezur-native scripts are planned for some languages; when they arrive, the romanized redirects become mandatory (custom-script titles cannot be typed on a Terran keyboard).
  • Homographs: the Etymology 1 / Etymology 2 pattern from the very first entry — built correctly from the get-go.
  • Inflected-form pages: deferred — the operator plans a conventional (non-LLM) bot for them.
  • Reconstructions: out of scope for now — and when they come, no separate namespace: ordinary Dictionary pages titled without the asterisk, the proto-language as a language section, the * shown in the headword line and in mentions only.
  • One translations box per heading, on the core meaning (operator, 2026-08-21): an English lemma carries one {{Translations}} box under each part-of-speech heading, translating that word's core meaning only. Extra senses that are English applications of the core — mother for a source, fire for a hearth — are definitions, not tables. A second table is warranted only where a second heading is: genuine homographs, split under Etymology 1 / Etymology 2 the way Dictionary:calf is. A language appears in a box only where its own core meaning matches; a word whose core is a neighbouring sense belongs on the English lemma that matches it, so Gaillean theste "offspring" is not a translation of child.
  • The gloss on a box must be unambiguous (operator, 2026-08-21): it is read on its own by someone choosing a translation, so a bare English word that English itself splits — hair (a head of hair, or one hair), stone, long, fly — will be translated wrong in both directions. Say which sense: gloss=the clear liquid that falls as rain and fills rivers, not gloss=water. Known offenders are tracked at Yezur Wiki:Ambiguous glosses.
  • Links are reciprocal (operator, 2026-08-21): if page A links to page B, page B links back to A. In practice: an English entry names a foreign form in its translations box, and that foreign entry's definition links back with {{d|word|YXS-EN}} — the shape used since Dictionary:geúrl. The same holds for etymologies and descendants, which have always been written as a matched pair.
  • Case-sensitive titles: done. The Dictionary and Dictionary-talk namespaces run case-sensitive (server-configured 2026-07-17), so polish/Polish-style pairs work.
  • Automatic pronunciation: done. PhoMo has been ported to Lua as Auphen (Module:Auphen); {{auphen|word|code}} returns an IPA estimation from a language's ruleset (per-language data, e.g. Module:Auphen/YBS-PJ). Divergences from the JavaScript original are logged at Module talk:Auphen.

Still open

  • Native scripts via Private Use Area fonts: feasible but deferred; the one-time setup needs interface rights (@font-face is not permitted in TemplateStyles, so the font must be wired up via MediaWiki:Common.css or the server). Input methods and font creation are the operator's sideline project.