Skip to content
#edge computing Dataset Open access

Sixth Buddhist Council Tipiṭaka (Chaṭṭhasaṅgītipiṭaka) — Unicode corpus

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

A searchable, cross-referenced Unicode corpus of the official Sixth Buddhist Council (Chaṭṭha Saṅgāyana) recension of the Pāḷi Tipiṭaka with its Aṭṭhakathā (commentaries) and Ṭīkā (subcommentaries) — 118 volumes, 89,512 paragraphs — converted from the legacy VZTimes-font edition published by the Ministry of Religious Affairs, Yangon (Pāḷi Series, romanised 2008 from the Myanmar version of 2001). Includes structural extraction anchored to the printed page (pages, paragraphs, suttas), 54,036 variant readings and 27,153 cross-references transcribed from the printed footnote apparatus with their sigla, and cross-layer paragraph linking. Counts are taken from site/ .json and site/reader/apparatus/ and can be re-derived from the deposit itself. New in 2.10.0: no text changed — a phrase is consecutive words, and two more files stop being downloaded whole. The corpus figures are identical to 2.9.0. Until now a phrase was counted as a substring of the paragraph's text, which counted tassa bhagavato inside etassa bhagavato (63 paragraphs) and refused dhammā"ti for dhammā ti, because the edition closes the quotative particle up against its word. A phrase now matches as consecutive tokens — each word as the single-word search matches it, the words following each other whatever punctuation the edition prints between them — and every phrase result line names the rule beside the diacritics mode. Counts move in both directions and the release note gives the table; a script in the deposit computes both rules over all 89,512 paragraphs. A position store that would decide adjacency without reading the text was measured and deliberately not built. Separately, a substring or *vaggo search no longer downloads the whole key list (12.5 MB) but one n-gram shard, and the section names (1.09 MB, read before every first search) are likewise read as one shard; the largest file a search fetches is now 514 KB, and a gate holds it there. Correctness is gated as before: every new assertion in check_search.js was made to fail on the previous build first. New in 2.9.0: no text changed — search matches diacritics exactly, and stops downloading the canon. The corpus figures are identical to 2.8.0. Until now every search key was folded (tassā stored as tassa), so a search for either word reported both: 36,644 occurrences where tassā alone has 4,322. The index now stores the printed tokens (NFC, lower case, the modern ṃ written as the edition's ṁ) and matches them by identity; 34,134 folded keys (5.3%) had been merging two or more printed forms. Diacritic folding is offered as a switch, "Ignore diacritics", off by default, and every result line names the mode that produced its count. At the same time the search stopped fetching whole volumes: counting one common word used to download 117 per-volume index files (about 40 MB compressed, 190 MB parsed); postings now live in 1,031 prefix shards and the paragraph text in 1,008 chunks fetched only for the rows drawn, so the same search moves 2.3 MB. One implementation (searchcore.js) serves both the search page and the reader's box. A performance gate records the before and after. The dictionary panel makes four round trips instead of six, and its files are now cached at the serving edge. Correctness is gated: the new assertions in check_search.js were made to fail on the previous build before the index was rebuilt, and folded-mode counts reproduce the previous release's exactly. New in 2.8.0: no text changed — the niggahita becomes the reader's choice. The corpus figures are identical to 2.7.1. The reader and the search page gain a toggle between the edition's ṁ and the modern ṃ (ṁ U+1E41 ↔ ṃ U+1E43, Ṁ ↔ Ṃ). The substitution is display-only and bijective: the stored corpus carries the edition's ṁ exclusively (verified: no ṃ occurs in any served volume, apparatus, link or index file), so the swap runs over what is on screen, never over the data, and switching back restores the display byte-exact. Links, anchors and citation identifiers are untouched; copied text follows the convention shown; search queries in either convention match the same passages, as they already did; a word clicked or typed in the dictionary panel is folded back to the edition's form before lookup. The choice persists per browser and is stated on the button itself. New in 2.7.1: ONE WORD of the text changes — the first release since 2.3.0 in which any reading moved, and it is stated here rather than left to be discovered. In 25VsmT01 the corrupt glyph in pathavīkasiṇādivaḍḍhaní was being restored as vaḍḍhanī; the maintainers settled it against the printed page as vaḍḍhane, and the superseded reading is kept in the register with the reason it fell — it was an inference from a pattern, never a witness, and the glyph occurs in exactly one place in the corpus. The change is one character for one character, so no offset, page break or paragraph boundary moved: 118 volumes, 89,512 paragraphs, 54,036 variant readings and 27,153 cross-references are unchanged from 2.3.0 through 2.7.0. Otherwise this release completes the work 2.7.0 left open. The maintainers' verdicts on the visual review sheet are applied: 41 register entries move to confirmed (49 of 68 now confirmed), and eleven suggested readings that were wrong in the register are corrected against the printed page, each with its superseded reading kept. Nine sites the review sheet had raised are WITHDRAWN rather than recorded as errata: the Unicode PDFs read correctly at every one, and the corrupt forms existed only in an unpublished working snapshot three weeks stale — a census over a stale artefact measures the artefact, not the edition. One further self-contradiction is corrected: the glyph register summarised a class as unanimous while one of its own entries had always dissented. New in 2.7.0: no text changed. The corpus figures are identical to 2.3.0 through 2.6.0. This release puts the errata register under scrutiny: entry E021 (10Ma02, printed p.247) is resolved as not an erratum — the printed x is the edition's own unnumbered position siglum, paired with the + eleven lines below, both keying one unnumbered foot-note; the corpus preserves both marks and the register records two superseded machine suggestions. A visual review instrument renders all 50 unresolved census-glyph sites beside their printed pages (clips included in this deposit); the maintainers' verdicts on 49 of them are recorded, confirming 37 corrections the served text already carries and leaving 12 for per-entry application. A companion measure sizes the unnumbered-siglum apparatus problem: of 2,535 paragraphs carrying a *, +, x or ( ) mark, 1,427 have no apparatus entry at all. New in 2.6.0: no text changed. The 118 volumes, 89,512 paragraphs, 54,036 variant readings and 27,153 cross-references are identical to 2.3.0 through 2.5.0; the paragraph figure was re-counted from the deposit, not carried forward. What changed: search results are ordered Pāḷi → Aṭṭhakathā → Ṭīkā with per-layer limits and one label idiom across both interfaces; the term index is served in prefix buckets, so the first search of a visit costs kilobytes instead of the whole 22 MB map; the dictionary panel gains its configurable defaults (CPED and PED open, a gear for the rest); the Digital Pāḷi Dictionary is refreshed to its 2026-07-28 release and its root-family, compound-family and idiom tables — never functional in any earlier deposit — now load with their content; 212 commentary vagga heads and 304 Dhammapada vatthu titles are reclassified against the printed page; and every page states this version, with its last-updated date, from a single source. New in 2.5.0: no text changed. The 118 volumes, 89,512 paragraphs, 54,036 variant readings and 27,153 cross-references are identical to 2.3.0 and 2.4.0; the paragraph figure was re-counted from the deposit, not carried forward. What changed is the reader: the printed footnote apparatus now reaches the page — markers restored where a bold lemma or a closing bracket broke them, and each printed page's notes drawn at the foot of that page — and both search interfaces answer multi-word and phrase queries, take a * wildcard, name the book of every result from the edition's own title stacks, list results Pāḷi → Aṭṭhakathā → Ṭīkā, and filter by layer. Every change is guarded by checks that drive the real reader and were made to fail on the build carrying the bug. New in 2.4.0: no text changed. The 118 volumes, 89,512 paragraphs, 54,036 variant readings and 27,153 cross-references are identical to 2.3.0, and a citation of either names the same text. What changed is that the dictionary stores moved out of the published site and are served from a separate origin, taking the site from 26,576 published files to 1,977. The stores remain inside this deposit — all 24,599 files — and the reader falls back to that copy when the serving origin cannot be reached, so an archived copy of this record still has working dictionaries with or without it. New in 2.3.0: four Khuddaka commentaries re-segmented against the printed page (20KhuA01, 21KhuA02, 23KhuA04, 24KhuA05 — 460 paragraphs become 3,607); a canon paragraph's commentary is now served as the whole printed range rather than its first paragraph; and the first checks that compare the corpus against the printed page rather than against itself. This is not the VRI / Igatpuri (CST) digital edition, which is a different recension with its own editorial layer. Where the two differ, this corpus follows the official printed edition. Browse the corpus at https://buddha-dhamma.net.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.