Module:Auphen/YBS-PJ: Difference between revisions
Add Pjany sound data |
Header comment: the frame loads this with require, not mw.loadData (flag 48b) |
||
| (7 intermediate revisions by one other user not shown) | |||
| Line 1: | Line 1: | ||
-- Module:Auphen/YBS-PJ -- sound data for the Pjany language (registry code | -- Module:Auphen/YBS-PJ -- sound data for the Pjany language (registry code | ||
-- YBS-PJ). Pure data, loaded by [[Module:Auphen/frame]] | -- YBS-PJ). Pure data, loaded by [[Module:Auphen/frame]]. | ||
-- ipacats : IPA-scope categories, used by pronunciation estimation | -- ipacats : IPA-scope categories, used by pronunciation estimation | ||
-- ipa_rules : glyph -> single IPA sound (applied before the ruleset runs) | -- ipa_rules : glyph -> single IPA sound (applied before the ruleset runs) | ||
| Line 20: | Line 20: | ||
}, | }, | ||
ipa_rules = [[ | ipa_rules = [[ | ||
ao/ɑ: | |||
ae/æ: | |||
éi/eɪ | |||
oú/oʊ | |||
au/ɐʊ | |||
aú/ɑu | |||
úi/ɥi | |||
ou/ɔə | |||
ui/əɪ | |||
qh/q͡χ | qh/q͡χ | ||
-/5 | |||
–/5 | |||
i/i | i/i | ||
p/p | p/p | ||
| Line 43: | Line 54: | ||
é/e | é/e | ||
ú/u | ú/u | ||
'/ | '/4 | ||
]], | ]], | ||
pronounce = [[ | pronounce = [[ | ||
| Line 80: | Line 91: | ||
[V]/%:8/_% | [V]/%:8/_% | ||
8[V]/ | 8[V]/ | ||
5/ | 5/~ | ||
:/ː | :/ː | ||
]], | ]], | ||
sets = {}, | sets = {}, | ||
} | } | ||
Latest revision as of 20:09, 28 July 2026
This is the documentation for Module:Auphen/YBS-PJ, the sound data for Pjany (registry code YBS-PJ) on the Yezur wiki. It is pure data and defines no functions: Module:Auphen/frame loads it whenever Template:Auphen is called as {{auphen|word|YBS-PJ}} and hands its tables to Module:Auphen, which does the work. The rule notation, and the engine's divergences from it, are documented on Module talk:Auphen.
Fields
| Field | Contents |
|---|---|
ipacats |
The category sets the ruleset works with, written [X]. They hold sounds, not letters, because they are read after the transposition. Pjany supplies ipacats only; Module:Auphen/frame falls back to a cats table where a language declares one for rulesets that operate on spelling.
|
ipa_rules |
The transposition from spelling to sound, applied once before pronounce runs. It is passed to the engine in pronunciation estimation only, so a named ruleset from sets would see the orthography untransposed.
|
pronounce |
The ruleset used for pronunciation estimation. |
sets |
Named rulesets for grammar and derivation. Empty for Pjany. |
Spelling and sound
The word is lower-cased before transposition, and the longest spelling that matches at a given point wins, so digraphs are read ahead of their component letters whatever order they stand in.
| Spelling | Sound | Spelling | Sound | |
|---|---|---|---|---|
a |
a |
ae |
æ:
| |
e |
ɛ |
ao |
ɑ:
| |
i |
i |
au |
ɐʊ
| |
o |
ɔ |
aú |
ɑu
| |
u |
ə |
éi |
eɪ
| |
é |
e |
ou |
ɔə
| |
ó |
o |
oú |
oʊ
| |
ú |
u |
ui |
əɪ
| |
úi |
ɥi
|
| Spelling | Sound | Spelling | Sound | |
|---|---|---|---|---|
b |
b |
n |
n
| |
d |
d |
p |
p
| |
g |
g |
q |
q
| |
h |
h |
qh |
q͡χ
| |
k |
k |
r |
ɹ
| |
l |
l |
s |
s
| |
m |
m |
t |
t
| |
z |
z
|
Two spellings carry no sound of their own. An apostrophe ' geminates the consonant before it; a hyphen -, or an en dash –, marks a seam within a compound. Both are carried through the ruleset as digits, 4 and 5, so that rules can see them: the 4 is dropped as soon as the doubling has applied, and the 5 becomes a ~ in the last rules, which Module:Auphen renders as a space. The digits 1, 2 and 3 do the same work for palatalisation and 8 for vowel length; none survive to the output. Length is written : throughout the data and rewritten to ː by the last rule.
An e that stands before another vowel is not read as ɛ where it opens the word or follows a consonant: there the ruleset takes it as a palatal element, described below.
Categories
Members are declared in order, and category-to-category mapping is positional, so the paired sets below line up member by member.
| Category | Members | Role |
|---|---|---|
[V] |
the vowel inventory, with the backed vowels of [B] appended |
Vowel environments. |
[C] |
the consonant inventory, including the palatals, lenited velars and uvulars that later rules derive | Consonant environments. |
[A] |
i e ɛ æ: æ a |
Front vowels. |
[K] |
p b k g t d |
Plosives. |
[R] |
r ʀ ɹ |
Rhotics. |
[Q] → [B] |
i ɛ e ə a æ ɐ → ɯ ʌ ɤ o̞ ɑ ɑ ɑ |
Backing next to a uvular. |
[H] → [J] |
s z l n t d p b h k g → ʃ ʒ ʎ ɲ tʃ dʒ ɸ β j x ɣ |
Palatalisation. |
[W] → [Y] |
g k q → ɣ x x |
Lenition after a lateral. |
[V] holds vowels proper: the offglides ɪ and ʊ and the glide ɥ that the diphthongs bring in are not members, so a rule keyed on [V] does not match them.
The pronunciation ruleset
pronounce applies the following, in this order.
| Process | Effect |
|---|---|
| Seam breathing | A seam standing between two vowels takes an inserted h.
|
| Gemination | A consonant before ' is doubled and the ' dropped. A glottal element is inserted between the two copies late in the ruleset, so t' surfaces as tˀt.
|
| Palatalisation | The palatal element surfaces as j word-initially and as a ʲ diacritic after ɹ or m, with ɹ hardening to r before the diacritic. Before a front vowel a preceding plosive instead takes an inserted j, unless that vowel is itself followed by q or qh. Otherwise the preceding consonant shifts to its palatal counterpart, [H] → [J]. After q the element is dropped and the q held aside as a marker until the backing rule has run, so the following vowel does not back.
|
| Uvular colouring | Beside q or q͡χ, front and central vowels back, [Q] → [B].
|
| Nasal assimilation | n becomes m before a labial, ŋ before a velar and ɴ before q.
|
| Rhotic backing | r and ɹ become ʀ beside q.
|
| Lenition after a lateral | g k q become ɣ x x after l or ʎ.
|
| L-vocalisation | l becomes ʊ before t d n s z and ə before p b m ɸ β, in both cases only when no consonant-plus-vowel follows. A word-final l after a consonant takes a preceding ə.
|
| R-vocalisation | A rhotic becomes ə after a front vowel when a consonant or the word end follows, and after ə when a consonant follows. An ə before a word-final rhotic opens to ɔ.
|
| Final adjustments | hə becomes x except word-finally; ɛə becomes ɛ:; ɥi becomes ɥʲ before a vowel; and two identical vowels merge into one long vowel.
|
A seam blocks adjacency, so rules that are meant to reach across one list it in their environments explicitly; the rest do not apply over it.
Named rulesets
sets is empty: Pjany has no derivation rulesets yet, and the three-argument form {{auphen|word|YBS-PJ|name}} therefore reports an unknown ruleset and files the page into Category:Auphen errors.
Four worked Pjany estimations are checked on every parse at Template:Auphen/testcases.
-- Module:Auphen/YBS-PJ -- sound data for the Pjany language (registry code
-- YBS-PJ). Pure data, loaded by [[Module:Auphen/frame]].
-- ipacats : IPA-scope categories, used by pronunciation estimation
-- ipa_rules : glyph -> single IPA sound (applied before the ruleset runs)
-- pronounce : the pronunciation-estimation ruleset
-- sets : named grammar/derivation rulesets (orthography scope)
return {
ipacats = {
['[A]'] = {'i','e','ɛ','æ:','æ','a'},
['[B]'] = {'ɯ','ʌ','ɤ','o̞','ɑ','ɑ','ɑ'},
['[C]'] = {'ɹ','s','z','l','n','t','d','m','p','b','h','ʔ','q','k','g','ʃ','ʒ','ʎ','ɲ','tʃ','dʒ','ɸ','β','j','k','x','ɣ','r','rʲ','mʲ','ʀ','ŋ','ɴ'},
['[H]'] = {'s','z','l','n','t','d','p','b','h','k','g'},
['[J]'] = {'ʃ','ʒ','ʎ','ɲ','tʃ','dʒ','ɸ','β','j','x','ɣ'},
['[K]'] = {'p','b','k','g','t','d'},
['[Q]'] = {'i','ɛ','e','ə','a','æ','ɐ'},
['[R]'] = {'r','ʀ','ɹ'},
['[V]'] = {'u','i','ɔ','ɛ','o','e','ə','ɑ','a','æ','ɐ','ɯ','ʌ','ɤ','o̞','ɑ','ɑ','ɑ'},
['[W]'] = {'g','k','q'},
['[Y]'] = {'ɣ','x','x'},
},
ipa_rules = [[
ao/ɑ:
ae/æ:
éi/eɪ
oú/oʊ
au/ɐʊ
aú/ɑu
úi/ɥi
ou/ɔə
ui/əɪ
qh/q͡χ
-/5
–/5
i/i
p/p
l/l
r/ɹ
z/z
a/a
q/q
h/h
o/ɔ
g/g
m/m
t/t
e/ɛ
s/s
d/d
k/k
u/ə
b/b
n/n
ó/o
é/e
ú/u
'/4
]],
pronounce = [[
5/5h/[V]_[V]
[C]/%%/_4
4/
ɛ/1/[C]_[V]|#_[V]|[C]5_[V]
1/j/#_
1/ʲ/ɹ_|m_|ɹ5_|m5_
1/2/_[A]/_[A]q|_[A]q͡χ
[K]/%j/_2|_52
2/1
ɹ/r/_ʲ
[H]/[J]/_1|_51
q1/3
q51/35
1/
2/
[Q]/[B]/q_|_q|_5q|q5_|q͡χ_|_q͡χ|_5q͡χ|q͡χ5_
3/q
n/m/_m|_b|_p|_5m|_5b|_5p
n/ŋ/_k|_g|_5k|_5g
n/ɴ/_q|_5q
r/ʀ/q_|_q|_5q|q5_
ɹ/ʀ/q_|_q|_5q|q5_
[W]/[Y]/l_|ʎ_|l5_|ʎ5_
l/ʊ/_t|_d|_n|_s|_z/_[C][V]
l/ə/_p|_b|_m|_ɸ|_β/_[C][V]
l/əl/[C]_#
[R]/ə/[A]_[C]|[A]_#|[A]5_[C]|[A]_5[C]|ə_[C]|ə5_[C]|ə_5[C]
ə/ɔ/_[R]#
[C]/%ˀ/_%
hə/x//_#
ɛə/ɛ:
ɥi/ɥʲ/_[V]
[V]/%:8/_%
8[V]/
5/~
:/ː
]],
sets = {},
}