Lexical Unit Boundary
The logical boundary of a lexical unit (word). It is used to mark the end of a word regardless of its physical layout on the support, allowing the system to reconstruct units split by fractures or line breaks.
Syntax
Any number of Unicode characters enclosed within ampersands &…&. Optionally, linguistic annotations can follow after a semicolon ;.
If the user wishes to easily identify a series of lexical units, the entire sequence can be enclosed within ampersands &…& with individual units separated by a forward slash /.
Overview
| Notation | Simplified AST | Paper Output |
|---|---|---|
&𝑆×𝑁;λₛ=λᵥ;…& | LexicalUnit([𝘜], LinguisticAnalysis{λₖ:λᵥ, …}) | 𝑆×𝑁 |
𝑆= any sign (Unicode character) ∉STOP_SET(cf. Technical Catalogue of Symbols) otherwise preceded by the escape character\.𝑁= number of characters (represented by char or digits).λₛ= label, which may also be in abbreviated form, for linguistic annotation.λᵥ= value of the linguistic annotation (plausibility scale).λₖ= canonical key for the linguistic annotation derived from the corresponding label.
Examples
| Notation | Paper Output |
|---|---|
| ab-c |
| ab c |
| abc defg hijk l |
| abcdefhijk |
EpiDoc Outputs
| Notation | EpiDoc (XML) |
|---|---|
| |
| |
| |
| |
Structural Analysis (Conceptual AST)
| Notation | AST |
|---|---|
| |
| |
| |
| |