Skip to content

Commit 1849079

Browse files
committed
readme
1 parent 7e60a50 commit 1849079

1 file changed

Lines changed: 0 additions & 95 deletions

File tree

README.md

Lines changed: 0 additions & 95 deletions
Original file line numberDiff line numberDiff line change
@@ -58,101 +58,6 @@ Preserving a complete semantic unit like a section, paragraph, sentence, etc., i
5858

5959
[Comparison of chunk size 200 with 1.5x overflow ratio: Chunkdown (left) / LangChain Markdown Splitter (right)](https://chunkdown.zirkelc.dev/?text=LSBbYGdlbmVyYXRlVGV4dGBdKC9kb2NzL2FpLXNkay1jb3JlL2dlbmVyYXRpbmctdGV4dCk6IEdlbmVyYXRlcyB0ZXh0IGFuZCBbdG9vbCBjYWxsc10oLi90b29scy1hbmQtdG9vbC1jYWxsaW5nKS4KICBUaGlzIGZ1bmN0aW9uIGlzIGlkZWFsIGZvciBub24taW50ZXJhY3RpdmUgdXNlIGNhc2VzIHN1Y2ggYXMgYXV0b21hdGlvbiB0YXNrcyB3aGVyZSB5b3UgbmVlZCB0byB3cml0ZSB0ZXh0IChlLmcuIGRyYWZ0aW5nIGVtYWlsIG9yIHN1bW1hcml6aW5nIHdlYiBwYWdlcykgYW5kIGZvciBhZ2VudHMgdGhhdCB1c2UgdG9vbHMuCi0gW2BzdHJlYW1UZXh0YF0oL2RvY3MvYWktc2RrLWNvcmUvZ2VuZXJhdGluZy10ZXh0KTogU3RyZWFtIHRleHQgYW5kIHRvb2wgY2FsbHMuCiAgWW91IGNhbiB1c2UgdGhlIGBzdHJlYW1UZXh0YCBmdW5jdGlvbiBmb3IgaW50ZXJhY3RpdmUgdXNlIGNhc2VzIHN1Y2ggYXMgW2NoYXQgYm90c10oL2RvY3MvYWktc2RrLXVpL2NoYXRib3QpIGFuZCBbY29udGVudCBzdHJlYW1pbmddKC9kb2NzL2FpLXNkay11aS9jb21wbGV0aW9uKS4KLSBbYGdlbmVyYXRlT2JqZWN0YF0oL2RvY3MvYWktc2RrLWNvcmUvZ2VuZXJhdGluZy1zdHJ1Y3R1cmVkLWRhdGEpOiBHZW5lcmF0ZXMgYSB0eXBlZCwgc3RydWN0dXJlZCBvYmplY3QgdGhhdCBtYXRjaGVzIGEgW1pvZF0oaHR0cHM6Ly96b2QuZGV2Lykgc2NoZW1hLgogIFlvdSBjYW4gdXNlIHRoaXMgZnVuY3Rpb24gdG8gZm9yY2UgdGhlIGxhbmd1YWdlIG1vZGVsIHRvIHJldHVybiBzdHJ1Y3R1cmVkIGRhdGEsIGUuZy4gZm9yIGluZm9ybWF0aW9uIGV4dHJhY3Rpb24sIHN5bnRoZXRpYyBkYXRhIGdlbmVyYXRpb24sIG9yIGNsYXNzaWZpY2F0aW9uIHRhc2tzLgotIFtgc3RyZWFtT2JqZWN0YF0oL2RvY3MvYWktc2RrLWNvcmUvZ2VuZXJhdGluZy1zdHJ1Y3R1cmVkLWRhdGEpOiBTdHJlYW0gYSBzdHJ1Y3R1cmVkIG9iamVjdCB0aGF0IG1hdGNoZXMgYSBab2Qgc2NoZW1hLgogIFlvdSBjYW4gdXNlIHRoaXMgZnVuY3Rpb24gdG8gW3N0cmVhbSBnZW5lcmF0ZWQgVUlzXSgvZG9jcy9haS1zZGstdWkvb2JqZWN0LWdlbmVyYXRpb24pLg%3D%3D&tab=aiSdk)
6060

61-
## How It Works
62-
63-
Chunkdown employs a sophisticated multi-layered approach that combines AST-based parsing with hierarchical processing to create semantically meaningful chunks while preserving markdown formatting.
64-
65-
### Algorithm Overview
66-
67-
Chunkdown uses a **hierarchical divide-and-conquer approach** that respects document structure:
68-
69-
1. **Structure Recognition**: Parse markdown into a tree where headings organize their related content into logical sections
70-
2. **Smart Chunking**: Keep complete sections together when possible, intelligently merge related sections to maximize space utilization
71-
3. **Graceful Degradation**: When sections are too large, progressively break them down using semantic boundaries (sentences, then phrases, then words)
72-
4. **Format Preservation**: Protect meaningful constructs like links and code from being split, maintaining markdown integrity throughout
73-
74-
### Step-by-Step Process
75-
76-
#### 1. AST Parsing and Hierarchical Transformation
77-
- Parse markdown using `mdast-util-from-markdown` with GitHub Flavored Markdown support
78-
- Transform flat AST into hierarchical sections using `createHierarchicalAST()`
79-
- **Key insight**: Headings become containers that hold their related content and nested subsections
80-
- Content size calculated from actual text content, not raw markdown characters
81-
82-
#### 2. Top-Down Section Processing
83-
- **Size evaluation**: Check if entire sections fit within `maxAllowedSize` (chunkSize × maxOverflowRatio)
84-
- **Keep together**: Sections within limits are preserved as single chunks to maintain semantic coherence
85-
- **Break down intelligently**: Large sections are decomposed using multiple optimization strategies
86-
87-
#### 3. Hierarchical Optimization Strategies
88-
89-
**Parent-Child Merging**:
90-
- Attempt to merge parent section (heading + immediate content) with child sections
91-
- Find consecutive children that fit together with parent within size limits
92-
- Create merged sections while preserving hierarchical relationships
93-
94-
**Sibling Section Merging**:
95-
- Group consecutive sibling sections at the same depth level
96-
- Create "orphaned sections" (no heading, depth 0) to combine related siblings
97-
- Maximize chunk utilization while respecting semantic boundaries
98-
99-
**Content Grouping**:
100-
- Within sections, group content items to maximize chunk space utilization
101-
- Use `flushCurrentItems()` pattern to accumulate content until size limits are reached
102-
- Handle both regular sections (with headings) and orphaned sections (content-only)
103-
104-
#### 4. Container-Specific Processing
105-
106-
**Lists**:
107-
- Process items individually while preserving list structure
108-
- Maintain correct numbering for ordered lists using `start` attribute
109-
- Group items when they fit within size limits
110-
111-
**Tables**:
112-
- Keep table headers with data rows when possible
113-
- Handle header-separator combinations for proper table formatting
114-
- Split by rows when table is too large
115-
116-
**Blockquotes**:
117-
- Treat as containers with recursive item processing
118-
- Preserve quote formatting across chunks
119-
120-
#### 5. Text-Level Fallback Mechanism
121-
122-
When hierarchical processing cannot reduce content to acceptable sizes:
123-
124-
**Protected Range Extraction**:
125-
- Re-parse content to extract position information for inline constructs
126-
- Protect links, images, inline code, emphasis, and other formatting from mid-construct splits
127-
- Create `ProtectedRange` objects that define no-split zones
128-
129-
**Priority-Based Boundary Detection**:
130-
- Extract semantic boundaries with priority hierarchy (lower number = higher priority):
131-
- Periods before newlines (priority 0)
132-
- Periods before uppercase letters (priority 1)
133-
- Question/exclamation marks (priority 2)
134-
- Safe periods (not abbreviations) (priority 3)
135-
- Colons and semicolons (priority 4)
136-
- Brackets and quotes (priority 5-8)
137-
- Line breaks, commas, dashes (priority 9-12)
138-
- Whitespace (priority 13, lowest)
139-
140-
**Recursive Boundary Splitting**:
141-
- Try boundaries in priority order (highest priority first)
142-
- Find optimal split position using middle-point strategy
143-
- Recursively process both parts with remaining lower-priority boundaries
144-
- Ensure forward progress by creating smaller parts than original
145-
146-
### Key Features in Action
147-
148-
**Content-Based Sizing**: Uses `toString()` to extract plain text from AST nodes, ensuring consistent semantic density regardless of markdown complexity.
149-
150-
**Controlled Overflow**: Allows chunks to exceed `chunkSize` up to `maxAllowedSize` to preserve complete semantic units like paragraphs or list items.
151-
152-
**Semantic Preservation**: Multi-level protection system prevents breaking meaningful constructs, from document structure (sections) down to inline elements (links, code spans).
153-
154-
**Structure Awareness**: Understands markdown document hierarchy and makes intelligent decisions about what content belongs together based on heading relationships.
155-
15661
## Usage
15762

15863
> [!NOTE]

0 commit comments

Comments
 (0)