Jump to content

Yezur Wiki:Dictionary structure: Difference between revisions

From Yezur Wiki
Alompan (talk | contribs)
Worked example: use a proto-language that exists (flag 45c)
Thiorosan (talk | contribs)
Cargo ruling 2026-07-30: Lexemes stays as the wiki's only Cargo table (deliberate exception); registry syncs to the infobox only
Line 53: Line 53:
== What the Yezur Dictionary adopts ==
== What the Yezur Dictionary adopts ==


The same skeleton, sized for a dictionary starting from zero on a Cargo-first wiki:
The same skeleton, sized for a dictionary starting from zero:


* '''Adopt as-is:''' one page per written form; one L2 section per Yezuri language; the fixed heading order and nesting (including ''Etymology 1 / Etymology 2'' for homographs, from the first entry); numbered senses; glosses in the wiki's prose language — in-world, Traditional English (see [[#The English frame|below]]).
* '''Adopt as-is:''' one page per written form; one L2 section per Yezuri language; the fixed heading order and nesting (including ''Etymology 1 / Etymology 2'' for homographs, from the first entry); numbered senses; glosses in the wiki's prose language — in-world, Traditional English (see [[#The English frame|below]]).
* '''Registry first — reusing the wiki's existing code system.''' [[Module:Languages]] is the registry: it maps the same codes the language infoboxes already carry (family-prefixed: <code>YAQ-GA</code> Gaillean, <code>YAQ-HE</code> Hertic; <code>YBS-LO</code> Lonish, <code>YBS-NI</code> Nichana, <code>YBS-PJ</code> Pjany, <code>YBS-UU</code> Uu; plus <code>YXS-EN</code> / <code>YXS-NE</code> for the two forms of English) to display name and Dictionary section heading. Every dictionary template takes the code as an argument, exactly as on Wiktionary. Keep the module in step with the <code>code</code> field of [[Template:Infobox language]] / the Languages Cargo table.
* '''Registry first — reusing the wiki's existing code system.''' [[Module:Languages]] is the registry: it maps the same codes the language infoboxes already carry (family-prefixed: <code>YAQ-GA</code> Gaillean, <code>YAQ-HE</code> Hertic; <code>YBS-LO</code> Lonish, <code>YBS-NI</code> Nichana, <code>YBS-PJ</code> Pjany, <code>YBS-UU</code> Uu; plus <code>YXS-EN</code> / <code>YXS-NE</code> for the two forms of English) to display name and Dictionary section heading. Every dictionary template takes the code as an argument, exactly as on Wiktionary. Keep the module in step with the <code>code</code> field of [[Template:Infobox language]].
* '''<code><nowiki>{{d|word|code}}</nowiki></code> — built''' ([[Template:D]]): links a word to its Dictionary entry, anchored to the language's section through the registry — <code><nowiki>{{d|dog|YXS-EN}}</nowiki></code> leads to the English section of ''Dictionary:dog''. Use it wherever prose refers to a word.
* '''<code><nowiki>{{d|word|code}}</nowiki></code> — built''' ([[Template:D]]): links a word to its Dictionary entry, anchored to the language's section through the registry — <code><nowiki>{{d|dog|YXS-EN}}</nowiki></code> leads to the English section of ''Dictionary:dog''. Use it wherever prose refers to a word.
* '''A minimal starter kit''' of further general-purpose templates (per-language headword templates can come later, once grammar exists to encode):
* '''A minimal starter kit''' of further general-purpose templates (per-language headword templates can come later, once grammar exists to encode):
Line 65: Line 65:
** <code><nowiki>{{inh}}</nowiki></code> / <code><nowiki>{{bor}}</nowiki></code> / <code><nowiki>{{cog}}</nowiki></code> — the etymology trio (still to build)
** <code><nowiki>{{inh}}</nowiki></code> / <code><nowiki>{{bor}}</nowiki></code> / <code><nowiki>{{cog}}</nowiki></code> — the etymology trio (still to build)
** <code><nowiki>{{ux|code|example}}</nowiki></code> — usage examples (still to build)
** <code><nowiki>{{ux|code|example}}</nowiki></code> — usage examples (still to build)
* '''Cargo from the first word.''' Where Wiktionary only categorizes, our <code><nowiki>{{head}}</nowiki></code> should also file each entry into a '''Lexemes''' Cargo table (word, language, part of speech, gloss, …). That is the interlingual payoff the Dictionary exists for: cognate sets, borrowing chains, "every Hertic noun", "all words inherited from Proto-Herta-Olaic" — one query each, self-maintaining.
* '''Cargo from the first word.''' Where Wiktionary only categorizes, our <code><nowiki>{{head}}</nowiki></code> also files each entry into the '''Lexemes''' Cargo table (word, language, part of speech, gloss, …). That is the interlingual payoff the Dictionary exists for: cognate sets, borrowing chains, "every Hertic noun", "all words inherited from Proto-Herta-Olaic" — one query each, self-maintaining. Since the encyclopedia retired its Cargo tables (2026-07), '''Lexemes is the wiki's only Cargo table''' — a deliberate exception, kept because the lexicon is the one place where scale outruns hand curation.
* '''Reconstructions in the same namespace''' (operator's call — no <code>Reconstruction:</code> analogue): a reconstructed form is an ordinary Dictionary page, its title the bare form '''without''' the asterisk; the proto-language is simply another language section. The '''headword line carries the leading <code>*</code>''' marking the form as unattested, as do mentions in running text — but never the page title.
* '''Reconstructions in the same namespace''' (operator's call — no <code>Reconstruction:</code> analogue): a reconstructed form is an ordinary Dictionary page, its title the bare form '''without''' the asterisk; the proto-language is simply another language section. The '''headword line carries the leading <code>*</code>''' marking the form as unattested, as do mentions in running text — but never the page title.
* '''No translation tables yet.''' With six languages and no lexicon, glosses carry the load; revisit the hub-language mechanism once entries exist in several languages for the same concepts.
* '''No translation tables yet.''' With six languages and no lexicon, glosses carry the load; revisit the hub-language mechanism once entries exist in several languages for the same concepts.

Revision as of 11:27, 30 July 2026

How the Dictionary namespace should be structured, modelled on Earth's Wiktionary (specifically the English Wiktionary, whose "Entry layout" policy is the most developed). This page is the out-of-frame blueprint; the Dictionary entries themselves are in-character. Nothing below is built yet — it documents the target so that words can be added in a structured, precise way from the very first entry.

How Wiktionary structures an entry

One page per written form. The page title is the exact spelling, case-sensitive (polish and Polish are different pages). Every language that uses that spelling shares the single page; a {{also|…}} hatnote at the very top cross-links near-identical spellings (case or diacritic variants).

One level-2 heading per language (==English==, ==Finnish==), in alphabetical order. Everything about the word in that language nests inside its section; languages never mix below level 2. A word's entry can be a two-line stub or a monograph, but it always sits in the same skeleton.

A fixed heading order inside each language section — only the headings that apply, but always in this order and nesting:

==Language==
===Alternative forms===
===Etymology===
===Pronunciation===        (IPA, audio, rhymes, hyphenation)
===Noun===                 (or Verb, Adjective, … — the part-of-speech heading)
Headword line
# Meaning 1
#* Quotations
====Usage notes====
====Synonyms==== / ====Antonyms==== / …
====Derived terms====
====Related terms====
====Descendants====
====Translations====       (English entries only)
===References===
===Anagrams===

Where one spelling covers several unrelated origins, ===Etymology 1===, ===Etymology 2===, … each carry their own part-of-speech sections, pushed one level deeper (====Noun====). Nesting is the key principle: a heading owns everything after it until a heading of the same level.

The headword line. Directly under each part-of-speech heading, a headword template renders the word in bold with its key grammar — language-specific where one exists ({{en-noun}}, which knows English plurals), the generic {{head|<code>|<pos>}} otherwise. It also auto-categorizes the page (Category:English nouns) — editors almost never hand-categorize.

Definitions are numbered wikitext list items (#), one sense per line. Context labels — {{lb|en|nautical|slang}} — prefix a sense and categorize it too. Example sentences ({{ux|…}}) and dated quotations (#*) nest under the sense they illustrate.

Everything is language-coded. The first argument of nearly every template is a language code (en, ang, gem-pro), resolved by one central registry module (their Module:languages) that knows each language's canonical name, script, and family. That single registry is what lets templates link, label, transliterate, and categorize automatically — it is the engine of the whole system.

Linking vs. mentioning. {{l|de|Hund}} links a term in a list (synonyms, derived terms); {{m|de|Hund}} mentions it in running prose (italicized, as in etymologies). Both target the right language's section of the right page.

Etymology templates display text and categorize in one stroke:

  • {{inh|en|ang|hund}} — inherited from an ancestor stage
  • {{bor|en|fr|hôtel}} — borrowed
  • {{der|…}} — derived, when the path is less direct
  • {{cog|de|Hund}} — a cognate in a sister language (display only)
  • {{doublet|en|canine}} — same-source pairs within the language

Non-English entries are glossed, not defined: a Finnish word gets a one-line English translation as its definition, and no Translations section. Translation tables ({{trans-top|gloss}} … {{t|de|Hund|m}} … {{trans-bottom}}, one collapsible table per sense) live only on English entries. English is the hub language — every pair of languages meets through it, which keeps the many-to-many translation problem linear.

Reconstructed terms are not mainspace. Proto-language words live in a dedicated Reconstruction: namespace (Reconstruction:Proto-Germanic/hundaz) and are always cited with a leading asterisk (*hundaz), the historical-linguistics convention for unattested forms.

Inflected forms get thin satellite pages ("hounds — plural of hound"), usually bot-created; the lemma page carries all real content.

What the Yezur Dictionary adopts

The same skeleton, sized for a dictionary starting from zero:

  • Adopt as-is: one page per written form; one L2 section per Yezuri language; the fixed heading order and nesting (including Etymology 1 / Etymology 2 for homographs, from the first entry); numbered senses; glosses in the wiki's prose language — in-world, Traditional English (see below).
  • Registry first — reusing the wiki's existing code system. Module:Languages is the registry: it maps the same codes the language infoboxes already carry (family-prefixed: YAQ-GA Gaillean, YAQ-HE Hertic; YBS-LO Lonish, YBS-NI Nichana, YBS-PJ Pjany, YBS-UU Uu; plus YXS-EN / YXS-NE for the two forms of English) to display name and Dictionary section heading. Every dictionary template takes the code as an argument, exactly as on Wiktionary. Keep the module in step with the code field of Template:Infobox language.
  • {{d|word|code}} — built (Template:D): links a word to its Dictionary entry, anchored to the language's section through the registry — {{d|dog|YXS-EN}} leads to the English section of Dictionary:dog. Use it wherever prose refers to a word.
  • A minimal starter kit of further general-purpose templates (per-language headword templates can come later, once grammar exists to encode):
    • {{head|code|pos}} — headword line + auto-categorization + a Lexemes row (built, Template:Head)
    • {{l|code|word}} and {{m|code|word}} — link and mention (built)
    • {{lb|code|label…}} — inline sense labels, italic and parenthesised before a sense (built, Template:Lb)
    • {{auphen|word|code}} — automatic pronunciation from the language's ruleset (built, Module:Auphen); plain {{IPA}} for hand-written IPA
    • {{inh}} / {{bor}} / {{cog}} — the etymology trio (still to build)
    • {{ux|code|example}} — usage examples (still to build)
  • Cargo from the first word. Where Wiktionary only categorizes, our {{head}} also files each entry into the Lexemes Cargo table (word, language, part of speech, gloss, …). That is the interlingual payoff the Dictionary exists for: cognate sets, borrowing chains, "every Hertic noun", "all words inherited from Proto-Herta-Olaic" — one query each, self-maintaining. Since the encyclopedia retired its Cargo tables (2026-07), Lexemes is the wiki's only Cargo table — a deliberate exception, kept because the lexicon is the one place where scale outruns hand curation.
  • Reconstructions in the same namespace (operator's call — no Reconstruction: analogue): a reconstructed form is an ordinary Dictionary page, its title the bare form without the asterisk; the proto-language is simply another language section. The headword line carries the leading * marking the form as unattested, as do mentions in running text — but never the page title.
  • No translation tables yet. With six languages and no lexicon, glosses carry the load; revisit the hub-language mechanism once entries exist in several languages for the same concepts.

The English frame

(Recorded from the operator, 2026-07-17; the in-world articles — Engil, the English language on Yezur — are still to be written.)

English is a language spoken on Yezur: a relatively small isolate spoken in the middle of the ocean, on the island nation of Engil. Its origins are nebulous and not well researched — which obscures the large amount of influence it has been under. It has two forms:

  • Traditional English (YXS-EN) — identical to Terran English in all ways except Yezuri semantic coinages. The wiki's prose is Traditional English, as most people on Engil speak it.
  • New English (YXS-NE) — the globalized form: the same language with roughly 3% of the vocabulary swapped for Anglified Yezuri loans or words of different etymology. In Terran terms, the replaced words fall into three main categories:
    1. words English loaned from non-European languages after the cutoff of roughly the 1500s (coffee, zebra, apartheidbeef is fully naturalized and stays; bouillon, c. 1600s, is a little too late and goes);
    2. words the world typically loans from English (computer, internet, hype);
    3. words with distinctly Terran origins (sandwich, rugby, saxophone, pasteurize).

The two forms are not treated as separate languages — comparable to how Wiktionary treats British and American English. One ==English== section serves both (both codes anchor there via the registry); a New English word such as tuku gets its own entry page, labelled as New English within its English section, and cross-referenced from its Traditional counterpart zebra — and vice versa. In article prose: "Zebras (New English: tuku) are equine animals…".

Settled decisions

  • Native-script page titles, with romanized forms as redirects. Yezur-native scripts are planned for some languages; when they arrive, the romanized redirects become mandatory (custom-script titles cannot be typed on a Terran keyboard).
  • Homographs: the Etymology 1 / Etymology 2 pattern from the very first entry — built correctly from the get-go.
  • Inflected-form pages: deferred — the operator plans a conventional (non-LLM) bot for them.
  • Reconstructions: out of scope for now — and when they come, no separate namespace: ordinary Dictionary pages titled without the asterisk, the proto-language as a language section, the * shown in the headword line and in mentions only.
  • Case-sensitive titles: done. The Dictionary and Dictionary-talk namespaces run case-sensitive (server-configured 2026-07-17), so polish/Polish-style pairs work.
  • Automatic pronunciation: done. PhoMo has been ported to Lua as Auphen (Module:Auphen); {{auphen|word|code}} returns an IPA estimation from a language's ruleset (per-language data, e.g. Module:Auphen/YBS-PJ). Divergences from the JavaScript original are logged at Module talk:Auphen.

Still open

  • Native scripts via Private Use Area fonts: feasible but deferred; the one-time setup needs interface rights (@font-face is not permitted in TemplateStyles, so the font must be wired up via MediaWiki:Common.css or the server). Input methods and font creation are the operator's sideline project.