You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: README.md
-95Lines changed: 0 additions & 95 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -58,101 +58,6 @@ Preserving a complete semantic unit like a section, paragraph, sentence, etc., i
58
58
59
59
[Comparison of chunk size 200 with 1.5x overflow ratio: Chunkdown (left) / LangChain Markdown Splitter (right)](https://chunkdown.zirkelc.dev/?text=LSBbYGdlbmVyYXRlVGV4dGBdKC9kb2NzL2FpLXNkay1jb3JlL2dlbmVyYXRpbmctdGV4dCk6IEdlbmVyYXRlcyB0ZXh0IGFuZCBbdG9vbCBjYWxsc10oLi90b29scy1hbmQtdG9vbC1jYWxsaW5nKS4KICBUaGlzIGZ1bmN0aW9uIGlzIGlkZWFsIGZvciBub24taW50ZXJhY3RpdmUgdXNlIGNhc2VzIHN1Y2ggYXMgYXV0b21hdGlvbiB0YXNrcyB3aGVyZSB5b3UgbmVlZCB0byB3cml0ZSB0ZXh0IChlLmcuIGRyYWZ0aW5nIGVtYWlsIG9yIHN1bW1hcml6aW5nIHdlYiBwYWdlcykgYW5kIGZvciBhZ2VudHMgdGhhdCB1c2UgdG9vbHMuCi0gW2BzdHJlYW1UZXh0YF0oL2RvY3MvYWktc2RrLWNvcmUvZ2VuZXJhdGluZy10ZXh0KTogU3RyZWFtIHRleHQgYW5kIHRvb2wgY2FsbHMuCiAgWW91IGNhbiB1c2UgdGhlIGBzdHJlYW1UZXh0YCBmdW5jdGlvbiBmb3IgaW50ZXJhY3RpdmUgdXNlIGNhc2VzIHN1Y2ggYXMgW2NoYXQgYm90c10oL2RvY3MvYWktc2RrLXVpL2NoYXRib3QpIGFuZCBbY29udGVudCBzdHJlYW1pbmddKC9kb2NzL2FpLXNkay11aS9jb21wbGV0aW9uKS4KLSBbYGdlbmVyYXRlT2JqZWN0YF0oL2RvY3MvYWktc2RrLWNvcmUvZ2VuZXJhdGluZy1zdHJ1Y3R1cmVkLWRhdGEpOiBHZW5lcmF0ZXMgYSB0eXBlZCwgc3RydWN0dXJlZCBvYmplY3QgdGhhdCBtYXRjaGVzIGEgW1pvZF0oaHR0cHM6Ly96b2QuZGV2Lykgc2NoZW1hLgogIFlvdSBjYW4gdXNlIHRoaXMgZnVuY3Rpb24gdG8gZm9yY2UgdGhlIGxhbmd1YWdlIG1vZGVsIHRvIHJldHVybiBzdHJ1Y3R1cmVkIGRhdGEsIGUuZy4gZm9yIGluZm9ybWF0aW9uIGV4dHJhY3Rpb24sIHN5bnRoZXRpYyBkYXRhIGdlbmVyYXRpb24sIG9yIGNsYXNzaWZpY2F0aW9uIHRhc2tzLgotIFtgc3RyZWFtT2JqZWN0YF0oL2RvY3MvYWktc2RrLWNvcmUvZ2VuZXJhdGluZy1zdHJ1Y3R1cmVkLWRhdGEpOiBTdHJlYW0gYSBzdHJ1Y3R1cmVkIG9iamVjdCB0aGF0IG1hdGNoZXMgYSBab2Qgc2NoZW1hLgogIFlvdSBjYW4gdXNlIHRoaXMgZnVuY3Rpb24gdG8gW3N0cmVhbSBnZW5lcmF0ZWQgVUlzXSgvZG9jcy9haS1zZGstdWkvb2JqZWN0LWdlbmVyYXRpb24pLg%3D%3D&tab=aiSdk)
60
60
61
-
## How It Works
62
-
63
-
Chunkdown employs a sophisticated multi-layered approach that combines AST-based parsing with hierarchical processing to create semantically meaningful chunks while preserving markdown formatting.
64
-
65
-
### Algorithm Overview
66
-
67
-
Chunkdown uses a **hierarchical divide-and-conquer approach** that respects document structure:
68
-
69
-
1.**Structure Recognition**: Parse markdown into a tree where headings organize their related content into logical sections
70
-
2.**Smart Chunking**: Keep complete sections together when possible, intelligently merge related sections to maximize space utilization
71
-
3.**Graceful Degradation**: When sections are too large, progressively break them down using semantic boundaries (sentences, then phrases, then words)
72
-
4.**Format Preservation**: Protect meaningful constructs like links and code from being split, maintaining markdown integrity throughout
73
-
74
-
### Step-by-Step Process
75
-
76
-
#### 1. AST Parsing and Hierarchical Transformation
77
-
- Parse markdown using `mdast-util-from-markdown` with GitHub Flavored Markdown support
78
-
- Transform flat AST into hierarchical sections using `createHierarchicalAST()`
79
-
-**Key insight**: Headings become containers that hold their related content and nested subsections
80
-
- Content size calculated from actual text content, not raw markdown characters
81
-
82
-
#### 2. Top-Down Section Processing
83
-
-**Size evaluation**: Check if entire sections fit within `maxAllowedSize` (chunkSize × maxOverflowRatio)
84
-
-**Keep together**: Sections within limits are preserved as single chunks to maintain semantic coherence
85
-
-**Break down intelligently**: Large sections are decomposed using multiple optimization strategies
86
-
87
-
#### 3. Hierarchical Optimization Strategies
88
-
89
-
**Parent-Child Merging**:
90
-
- Attempt to merge parent section (heading + immediate content) with child sections
91
-
- Find consecutive children that fit together with parent within size limits
92
-
- Create merged sections while preserving hierarchical relationships
93
-
94
-
**Sibling Section Merging**:
95
-
- Group consecutive sibling sections at the same depth level
96
-
- Create "orphaned sections" (no heading, depth 0) to combine related siblings
97
-
- Maximize chunk utilization while respecting semantic boundaries
98
-
99
-
**Content Grouping**:
100
-
- Within sections, group content items to maximize chunk space utilization
101
-
- Use `flushCurrentItems()` pattern to accumulate content until size limits are reached
102
-
- Handle both regular sections (with headings) and orphaned sections (content-only)
103
-
104
-
#### 4. Container-Specific Processing
105
-
106
-
**Lists**:
107
-
- Process items individually while preserving list structure
108
-
- Maintain correct numbering for ordered lists using `start` attribute
109
-
- Group items when they fit within size limits
110
-
111
-
**Tables**:
112
-
- Keep table headers with data rows when possible
113
-
- Handle header-separator combinations for proper table formatting
114
-
- Split by rows when table is too large
115
-
116
-
**Blockquotes**:
117
-
- Treat as containers with recursive item processing
118
-
- Preserve quote formatting across chunks
119
-
120
-
#### 5. Text-Level Fallback Mechanism
121
-
122
-
When hierarchical processing cannot reduce content to acceptable sizes:
123
-
124
-
**Protected Range Extraction**:
125
-
- Re-parse content to extract position information for inline constructs
126
-
- Protect links, images, inline code, emphasis, and other formatting from mid-construct splits
127
-
- Create `ProtectedRange` objects that define no-split zones
128
-
129
-
**Priority-Based Boundary Detection**:
130
-
- Extract semantic boundaries with priority hierarchy (lower number = higher priority):
131
-
- Periods before newlines (priority 0)
132
-
- Periods before uppercase letters (priority 1)
133
-
- Question/exclamation marks (priority 2)
134
-
- Safe periods (not abbreviations) (priority 3)
135
-
- Colons and semicolons (priority 4)
136
-
- Brackets and quotes (priority 5-8)
137
-
- Line breaks, commas, dashes (priority 9-12)
138
-
- Whitespace (priority 13, lowest)
139
-
140
-
**Recursive Boundary Splitting**:
141
-
- Try boundaries in priority order (highest priority first)
142
-
- Find optimal split position using middle-point strategy
143
-
- Recursively process both parts with remaining lower-priority boundaries
144
-
- Ensure forward progress by creating smaller parts than original
145
-
146
-
### Key Features in Action
147
-
148
-
**Content-Based Sizing**: Uses `toString()` to extract plain text from AST nodes, ensuring consistent semantic density regardless of markdown complexity.
149
-
150
-
**Controlled Overflow**: Allows chunks to exceed `chunkSize` up to `maxAllowedSize` to preserve complete semantic units like paragraphs or list items.
151
-
152
-
**Semantic Preservation**: Multi-level protection system prevents breaking meaningful constructs, from document structure (sections) down to inline elements (links, code spans).
153
-
154
-
**Structure Awareness**: Understands markdown document hierarchy and makes intelligent decisions about what content belongs together based on heading relationships.
0 commit comments