difflib

    A faithful MoonBit port of Python's difflib: SequenceMatcher, Differ, ndiff, unified/context diffs, HtmlDiff and get_close_matches.

    diff
    difflib
    sequence-matcher
    unified-diff
    ndiff
    Download zip
    Author
    Version
    0.2.0
    License
    Apache-2.0
    Last updated
    yesterday
    Downloads
    8

    #bobzhang/difflib

    A MoonBit port of Python's difflib: SequenceMatcher, Differ/ndiff, restore, unified_diff, context_diff, diff_bytes, HtmlDiff and get_close_matches.

    It is ported from CPython's Lib/difflib.py (main branch, commit 763b6edb0ec9) and aims for identical output. The test suite includes:

    • ports of CPython's Lib/test/test_difflib.py and the module's doctests;
    • a byte-for-byte comparison with CPython's test_difflib_expect.html;
    • a differential corpus of 1456 randomized cases generated by running the pinned CPython source (tools/gen_corpus.py). It covers opcodes, grouped opcodes, ratios, junk and autojunk handling, stateful set_seq1/set_seq2 call sequences, ndiff, restore, unified/context/colored diffs, diff_bytes, HTML tables and files, close matches and error messages. Inputs include tabs, Unicode whitespace, astral characters, lone surrogates and literal NUL/SOH characters (which collide with HtmlDiff's internal markers, exactly as in Python).

    #Python → MoonBit

    PythonMoonBit
    SequenceMatcher(isjunk, a, b, autojunk)SequenceMatcher::new(isjunk?, a?, b?, autojunk?) (generic over T : Hash + Eq)
    Match(a, b, size)Match { a, b, size }
    opcode tuples (tag, i1, i2, j1, j2)Opcode { tag : Tag, i1, i2, j1, j2 }
    get_close_matches(word, poss, n, cutoff)get_close_matches(word, poss, n?, cutoff?, autojunk?)
    Differ(linejunk, charjunk).compare(a, b)Differ::new(linejunk?, charjunk?).compare(a, b)
    ndiff, restorendiff, restore
    unified_diff, context_diffunified_diff, context_diff
    diff_bytes(unified_diff, ...)diff_bytes(Unified, ...) / diff_bytes(Context, ...)
    HtmlDiff(...).make_file / make_tableHtmlDiff::new(...).make_file / make_table
    IS_LINE_JUNK, IS_CHARACTER_JUNKis_line_junk, is_character_junk
    ValueErrorDiffError::ValueError(message) (same messages)

    To compare strings character by character, pass s.to_array(). Python indexes strings by code point, and so does Array[Char].

    Functions that are generators in Python return a lazy, single-pass Iter: Differ::compare, ndiff, restore, unified_diff, context_diff, diff_bytes and SequenceMatcher::get_grouped_opcodes. As in Python, no work (including junk callbacks) happens until the first element is requested, and lines are produced as they are consumed. Use .join(""), .to_array() or a for loop to consume them. Functions that return lists or strings in Python (get_opcodes, get_matching_blocks, get_close_matches, HtmlDiff) return an Array or a String.

    #Examples

    #SequenceMatcher

    ///|
    test "sequence matcher" {
    let s = @difflib.SequenceMatcher::new(
    a="qabxcd".to_array(),
    b="abycdf".to_array(),
    )
    inspect(
    s
    .get_opcodes()
    .map(op => "\{op.tag} a[\{op.i1}:\{op.i2}] b[\{op.j1}:\{op.j2}]")
    .join("\n"),
    content=(
    #|delete a[0:1] b[0:0]
    #|equal a[1:3] b[0:2]
    #|replace a[3:4] b[2:3]
    #|equal a[4:6] b[3:5]
    #|insert a[6:6] b[5:6]
    ),
    )
    inspect(s.ratio(), content="0.6666666666666666")
    }

    #Unified diff

    ///|
    test "unified diff" {
    let before = ["one\n", "two\n", "three\n", "four\n"]
    let after = ["zero\n", "one\n", "tree\n", "four\n"]
    inspect(
    @difflib.unified_diff(
    before,
    after,
    fromfile="before.txt",
    tofile="after.txt",
    ).join(""),
    content=(
    #|--- before.txt
    #|+++ after.txt
    #|@@ -1,4 +1,4 @@
    #|+zero
    #| one
    #|-two
    #|-three
    #|+tree
    #| four
    #|
    ),
    )
    }

    #ndiff and restore

    ///|
    test "ndiff" {
    let a = ["one\n", "two\n", "three\n"]
    let b = ["ore\n", "tree\n", "emu\n"]
    let delta = @difflib.ndiff(a, b).to_array()
    inspect(
    delta.join(""),
    content=(
    #|- one
    #|? ^
    #|+ ore
    #|? ^
    #|- two
    #|- three
    #|? -
    #|+ tree
    #|+ emu
    #|
    ),
    )
    assert_eq(@difflib.restore(delta.iter(), 1).to_array(), a)
    assert_eq(@difflib.restore(delta.iter(), 2).to_array(), b)
    }

    #Close matches

    ///|
    test "close matches" {
    debug_inspect(
    @difflib.get_close_matches("appel", ["ape", "apple", "peach", "puppy"]),
    content=(
    #|["apple", "ape"]
    ),
    )
    }

    #HTML side-by-side diff

    ///|
    test "html diff" {
    let html = @difflib.HtmlDiff::new(wrapcolumn=40).make_file(
    ["a\n", "b\n"],
    ["a\n", "c\n"],
    fromdesc="old",
    todesc="new",
    )
    assert_true(html.has_prefix("\n<!DOCTYPE html>"))
    }

    #Differences from Python

    • Python's restore checks which when its generator is first advanced; here the ValueError is raised immediately.
    • Python's get_grouped_opcodes overwrites the first and last entries of the opcode list cached by get_opcodes(). This port works on a copy and leaves the cache unchanged.
    • charjunk defaults to is_character_junk in ndiff and HtmlDiff. To get Python's charjunk=None, pass charjunk=_ => false, which behaves the same.
    • unified_diff(color=true) always emits ANSI escapes. It does not check NO_COLOR, FORCE_COLOR or whether the output is a terminal.
    • diff_bytes takes a DiffFormat (Unified or Context) rather than a diff function. Each byte becomes one character in the range U+0000–U+00FF and is encoded back afterwards. This is lossless and gives the same output as Python's ascii/surrogateescape round trip.
    • HtmlDiff::make_file(charset=...) emulates encode(charset,"xmlcharrefreplace") for the ASCII, Latin-1 and UTF-8/16/32 codecs and all of their Python aliases (US-ASCII, ISO-8859-1, latin1, utf8, ...). It assumes that any other charset name can encode all of Unicode, where Python would use that codec or raise LookupError.
    • Like Python's class attribute HtmlDiff._default_prefix, the anchor prefix counter is global and increases on every make_table/make_file call. HtmlDiff::reset_prefix_counter() resets it, for example to get reproducible output in tests.
    • Negative context sizes (n in get_grouped_opcodes, unified_diff, context_diff; numlines in HtmlDiff) and a negative wrapcolumn panic. Depending on the input, Python can return meaningless hunks or raise IndexError, ZeroDivisionError or RecursionError for these.
    • MoonBit strings are UTF-16, so a high surrogate directly followed by a low surrogate is the astral character they encode. Python can keep them as two separate code points; lone surrogates otherwise behave as in Python.
    • Python's runtime type checks (_check_types) are unnecessary under MoonBit's static types and were dropped.

    #Development

    Install the MoonBit toolchain, then run:

    moon test # one backend; add --target js, native, wasm or all

    The fixtures are generated from the pinned CPython sources in .repos/ (ignored by git). CI checks that they are up to date:

    git init .repos/cpython && cd .repos/cpython git remote add origin https://github.com/python/cpython git sparse-checkout set --no-cone /Lib/difflib.py /Lib/test/test_difflib.py /Lib/test/test_difflib_expect.html git fetch --depth 1 --filter=blob:none origin 763b6edb0ec959bdfb78f098b76829365de36ca5 git checkout FETCH_HEAD && cd ../.. python3 tools/gen_corpus.py > corpus/corpus_fixtures_test.mbt python3 tools/gen_html_fixtures.py .repos/cpython/Lib tools > html_fixtures_test.mbt python3 tools/gen_templates.py .repos/cpython/Lib > html_templates.mbt python3 tools/gen_charset.py > charset.mbt moon fmt # generated files are committed in formatted form

    #License

    Apache-2.0 (see LICENSE). This is a derivative of CPython's difflib. The derived portions retain the Python Software Foundation's copyright notice and license: see NOTICE and LICENSE-PSF.

    DiffError

    pub(all) suberror DiffError {
    ValueError(String)
    } derive(Eq,
    Debug
    )

    Errors raised for invalid arguments, mirroring Python's ValueError.

    DiffFormat

    pub(all) enum DiffFormat {
    Unified
    Context
    } derive(Eq,
    Debug
    )

    The output format used by [diff_bytes].

    Differ

    pub struct Differ {
    // private fields
    }

    Produces human-readable deltas from sequences of lines of text.

    Each line of a Differ delta begins with a two-letter code:

    • "- ": line unique to sequence 1
    • "+ ": line unique to sequence 2
    • " ": line common to both sequences
    • "? ": line not present in either input sequence, guiding the eye to intraline differences

    Example

    test {
    let text1 = [
    " 1. Beautiful is better than ugly.\n", " 2. Explicit is better than implicit.\n",
    " 3. Simple is better than complex.\n", " 4. Complex is better than complicated.\n",
    ]
    let text2 = [
    " 1. Beautiful is better than ugly.\n", " 3. Simple is better than complex.\n",
    " 4. Complicated is better than complex.\n", " 5. Flat is better than nested.\n",
    ]
    inspect(
    @difflib.Differ::new().compare(text1, text2).join(""),
    content=(
    #| 1. Beautiful is better than ugly.
    #|- 2. Explicit is better than implicit.
    #|- 3. Simple is better than complex.
    #|+ 3. Simple is better than complex.
    #|? ++
    #|- 4. Complex is better than complicated.
    #|? ^ ---- ^
    #|+ 4. Complicated is better than complex.
    #|? ++++ ^ ^
    #|+ 5. Flat is better than nested.
    #|
    ),
    )
    }

    Differ::compare

    fn Differ::compare(self : Differ, a : Array[String], b : Array[String]) -> Iter[String]

    Compares two sequences of lines and returns the resulting delta.

    Each line should end with a newline; the delta lines then also end with newlines.

    Like Python's generator, the delta is produced lazily: no work is done until the first line is requested, and the returned iterator can be consumed only once.

    Example

    test {
    let delta = @difflib.Differ::new().compare(["one\n", "two\n", "three\n"], [
    "ore\n", "tree\n", "emu\n",
    ])
    inspect(
    delta.join(""),
    content=(
    #|- one
    #|? ^
    #|+ ore
    #|? ^
    #|- two
    #|- three
    #|? -
    #|+ tree
    #|+ emu
    #|
    ),
    )
    }

    Differ::new

    fn Differ::new(linejunk? : (String) -> Bool, charjunk? : (Char) -> Bool, autojunk? : Bool) -> Differ

    Constructs a text differencer with optional junk filters.

    • linejunk: returns true iff a line is junk (e.g. [is_line_junk]). It is recommended to leave this unset.
    • charjunk: returns true iff a character is junk (e.g. [is_character_junk]).
    • autojunk: the automatic junk heuristic of [SequenceMatcher].

    HtmlDiff

    pub struct HtmlDiff {
    // private fields
    }

    Produces an HTML side by side comparison with change highlights.

    The table can be generated in either full or contextual difference mode, either as a complete HTML file ([HtmlDiff::make_file]) or as a table only ([HtmlDiff::make_table]).

    Example

    test {
    @difflib.HtmlDiff::reset_prefix_counter()
    let table = @difflib.HtmlDiff::new().make_table(["a\n", "b\n"], ["a\n", "c\n"])
    assert_true(table.contains("<span class=\"diff_sub\">b</span>"))
    assert_true(table.contains("<span class=\"diff_add\">c</span>"))
    }

    HtmlDiff::make_file

    fn HtmlDiff::make_file(self : HtmlDiff, fromlines : Array[String], tolines : Array[String], fromdesc? : String, todesc? : String, context? : Bool, numlines? : Int, charset? : String) -> String

    Returns a complete HTML file containing a side by side comparison table with change highlights.

    • fromdesc, todesc: column header strings.
    • context: show contextual differences (default false: full).
    • numlines: the number of context lines when context is set; otherwise how many lines before a change the "next" anchors are placed. Panics if negative.
    • charset: the document charset. Characters that the charset cannot encode are replaced by XML character references (&#NNN;). The ascii, latin-1 and UTF-8/16/32 codecs and their Python aliases are recognised; any other name is assumed to cover all of Unicode.

    HtmlDiff::make_table

    fn HtmlDiff::make_table(self : HtmlDiff, fromlines : Array[String], tolines : Array[String], fromdesc? : String, todesc? : String, context? : Bool, numlines? : Int) -> String

    Returns an HTML table of a side by side comparison with change highlights.

    The arguments have the same meaning as for [HtmlDiff::make_file].

    HtmlDiff::new

    fn HtmlDiff::new(tabsize? : Int, wrapcolumn? : Int, linejunk? : (String) -> Bool, charjunk? : (Char) -> Bool, autojunk? : Bool) -> HtmlDiff

    Constructs an HtmlDiff.

    • tabsize: tab stop spacing (default 8).
    • wrapcolumn: column where lines are broken and wrapped; by default (or when 0) lines are not wrapped. Panics if negative (Python recurses without bound).
    • linejunk, charjunk, autojunk: passed to [ndiff] (charjunk defaults to [is_character_junk]).

    HtmlDiff::reset_prefix_counter

    fn HtmlDiff::reset_prefix_counter(value? : Int) -> Unit

    Resets the global anchor-prefix counter shared by all [HtmlDiff] instances (Python's HtmlDiff._default_prefix = value). Each call to [HtmlDiff::make_table] or [HtmlDiff::make_file] uses the current value and then increments it.

    Match

    pub(all) struct Match {
    a : Int
    b : Int
    size : Int
    } derive(Compare, Eq, ToJson,
    Debug
    )

    A matching block: a[a:a+size] == b[b:b+size].

    Mirrors Python's difflib.Match named tuple.

    Opcode

    pub(all) struct Opcode {
    tag : Tag
    i1 : Int
    i2 : Int
    j1 : Int
    j2 : Int
    } derive(Eq, ToJson,
    Debug
    )

    A 5-tuple (tag, i1, i2, j1, j2) describing how to turn a[i1:i2] into b[j1:j2].

    SequenceMatcher

    pub struct SequenceMatcher[T] {
    // private fields
    }

    Compares pairs of sequences of any hashable element type.

    This is a port of Python's difflib.SequenceMatcher (Ratcliff/Obershelp "gestalt pattern matching" with junk handling and the "popular element" autojunk heuristic). It finds the longest contiguous junk-free matching subsequence and recurses on the pieces to its left and right.

    Strings are compared character by character by converting them with String::to_array().

    Example

    test {
    let s = @difflib.SequenceMatcher::new(
    isjunk=c => c == ' ',
    a="private Thread currentThread;".to_array(),
    b="private volatile Thread currentThread;".to_array(),
    )
    inspect((s.ratio() * 100).round() / 100, content="0.87")
    debug_inspect(
    s.get_matching_blocks(),
    content=(
    #|[
    #| { a: 0, b: 0, size: 8 },
    #| { a: 8, b: 17, size: 21 },
    #| { a: 29, b: 38, size: 0 },
    #|]
    ),
    )
    }

    SequenceMatcher::bjunk

    fn[T] SequenceMatcher::bjunk(self : SequenceMatcher[T]) -> Array[T]

    The elements of b for which isjunk returned true.

    SequenceMatcher::bpopular

    fn[T] SequenceMatcher::bpopular(self : SequenceMatcher[T]) -> Array[T]

    The non-junk elements of b treated as junk by the autojunk heuristic.

    SequenceMatcher::find_longest_match

    fn[T : Hash + Eq] SequenceMatcher::find_longest_match(self : SequenceMatcher[T], alo? : Int, ahi? : Int, blo? : Int, bhi? : Int) -> Match

    Finds the longest matching block in a[alo:ahi] and b[blo:bhi].

    Without junk, returns Match(i, j, k) such that a[i:i+k] == b[j:j+k] with k maximal; among maximal blocks the one starting earliest in a, then earliest in b, is returned. With junk, the longest junk-free block is found first and then extended by matching junk on both sides.

    If no blocks match, returns Match(alo, blo, 0).

    Example

    test {
    let s = @difflib.SequenceMatcher::new(
    a=" abcd".to_array(),
    b="abcd abcd".to_array(),
    )
    debug_inspect(s.find_longest_match(), content="{ a: 0, b: 4, size: 5 }")
    let s = @difflib.SequenceMatcher::new(
    isjunk=c => c == ' ',
    a=" abcd".to_array(),
    b="abcd abcd".to_array(),
    )
    debug_inspect(s.find_longest_match(), content="{ a: 1, b: 0, size: 4 }")
    }

    SequenceMatcher::get_grouped_opcodes

    fn[T : Hash + Eq] SequenceMatcher::get_grouped_opcodes(self : SequenceMatcher[T], n? : Int) -> Iter[Array[Opcode]]

    Isolates change clusters by eliminating ranges with no changes.

    Returns groups of opcodes with up to n lines of context each. Like Python's generator, the groups are computed when the returned iterator is first advanced, and it can be consumed only once.

    Unlike Python, the cached result of get_opcodes() is not mutated.

    Example

    test {
    let a = Array::makei(39, i => (i + 1).to_string())
    let b = a.copy()
    b.insert(8, "i") // Make an insertion
    b[20] = b[20] + "x" // Make a replacement
    for _ in 0..<5 {
    b.remove(23) |> ignore // Make a deletion
    }
    b[30] = b[30] + "y" // Make another replacement
    let groups = @difflib.SequenceMatcher::new(a~, b~).get_grouped_opcodes()
    inspect(
    groups
    .map(g => {
    g.map(op => "\{op.tag} \{op.i1} \{op.i2} \{op.j1} \{op.j2}").join(", ")
    })
    .join("\n"),
    content=(
    #|equal 5 8 5 8, insert 8 8 8 9, equal 8 11 9 12
    #|equal 16 19 17 20, replace 19 20 20 21, equal 20 22 21 23, delete 22 27 23 23, equal 27 30 23 26
    #|equal 31 34 27 30, replace 34 35 30 31, equal 35 38 31 34
    ),
    )
    }

    Panics

    Panics if n is negative (Python returns meaningless ranges).

    SequenceMatcher::get_matching_blocks

    fn[T : Hash + Eq] SequenceMatcher::get_matching_blocks(self : SequenceMatcher[T]) -> Array[Match]

    Returns the list of matching blocks (i, j, n) with a[i:i+n] == b[j:j+n], monotonically increasing in i and j. Adjacent blocks are collapsed, and the last block is always the dummy (len(a), len(b), 0).

    Example

    test {
    let s = @difflib.SequenceMatcher::new(
    a="abxcd".to_array(),
    b="abcd".to_array(),
    )
    debug_inspect(
    s.get_matching_blocks(),
    content=(
    #|[
    #| { a: 0, b: 0, size: 2 },
    #| { a: 3, b: 2, size: 2 },
    #| { a: 5, b: 4, size: 0 },
    #|]
    ),
    )
    }

    SequenceMatcher::get_opcodes

    fn[T : Hash + Eq] SequenceMatcher::get_opcodes(self : SequenceMatcher[T]) -> Array[Opcode]

    Returns the list of opcodes describing how to turn a into b.

    The first opcode has i1 == j1 == 0, and each subsequent opcode starts where the previous one ended.

    Example

    test {
    let a = "qabxcd"
    let b = "abycdf"
    let s = @difflib.SequenceMatcher::new(a=a.to_array(), b=b.to_array())
    let lines = s
    .get_opcodes()
    .map(op => "\{op.tag} a[\{op.i1}:\{op.i2}] b[\{op.j1}:\{op.j2}]")
    inspect(
    lines.join("\n"),
    content=(
    #|delete a[0:1] b[0:0]
    #|equal a[1:3] b[0:2]
    #|replace a[3:4] b[2:3]
    #|equal a[4:6] b[3:5]
    #|insert a[6:6] b[5:6]
    ),
    )
    }

    SequenceMatcher::new

    fn[T : Hash + Eq] SequenceMatcher::new(isjunk? : (T) -> Bool, a? : Array[T], b? : Array[T], autojunk? : Bool) -> SequenceMatcher[T]

    Constructs a SequenceMatcher.

    • isjunk: returns true iff an element of b is junk. Junk elements never start a match, but may extend one.
    • a, b: the sequences to compare (default empty).
    • autojunk: enables the "popular element" heuristic: when b has at least 200 elements, elements occurring more than 1% + 1 times are treated as junk.

    SequenceMatcher::quick_ratio

    fn[T : Hash + Eq] SequenceMatcher::quick_ratio(self : SequenceMatcher[T]) -> Double

    Returns an upper bound on ratio() relatively quickly, by treating both sequences as multisets.

    SequenceMatcher::ratio

    fn[T : Hash + Eq] SequenceMatcher::ratio(self : SequenceMatcher[T]) -> Double

    Returns a measure of the sequences' similarity as a float in [0, 1].

    Where T is the total number of elements in both sequences and M is the number of matches, this is 2.0 * M / T.

    Example

    test {
    let s = @difflib.SequenceMatcher::new(
    a="abcd".to_array(),
    b="bcde".to_array(),
    )
    inspect(s.ratio(), content="0.75")
    inspect(s.quick_ratio(), content="0.75")
    inspect(s.real_quick_ratio(), content="1")
    }

    SequenceMatcher::real_quick_ratio

    fn[T] SequenceMatcher::real_quick_ratio(self : SequenceMatcher[T]) -> Double

    Returns an upper bound on ratio() very quickly, using only the lengths.

    SequenceMatcher::set_seq1

    fn[T] SequenceMatcher::set_seq1(self : SequenceMatcher[T], a : Array[T]) -> Unit

    Sets the first sequence to be compared; the second is unchanged.

    Information about the second sequence is cached, so to compare one sequence against many, use set_seq2 once and set_seq1 repeatedly.

    SequenceMatcher::set_seq2

    fn[T : Hash + Eq] SequenceMatcher::set_seq2(self : SequenceMatcher[T], b : Array[T]) -> Unit

    Sets the second sequence to be compared; the first is unchanged.

    SequenceMatcher::set_seqs

    fn[T : Hash + Eq] SequenceMatcher::set_seqs(self : SequenceMatcher[T], a : Array[T], b : Array[T]) -> Unit

    Sets both sequences to be compared.

    Tag

    pub(all) enum Tag {
    Replace
    Delete
    Insert
    Equal
    } derive(Eq, ToJson,
    Debug
    )

    The kind of edit an [Opcode] describes.
    impl Show for Tag

    Tag::name

    fn Tag::name(self : Tag) -> String

    Returns the tag name used by Python ("replace", "delete", "insert", "equal").

    context_diff

    fn context_diff(a : Array[String], b : Array[String], fromfile? : String, tofile? : String, fromfiledate? : String, tofiledate? : String, n? : Int, lineterm? : String, autojunk? : Bool) -> Iter[String]

    Compares two sequences of lines and returns the delta as a context diff.

    The parameters have the same meaning as for [unified_diff], and the lines are likewise generated lazily.

    Example

    test {
    let diff = @difflib.context_diff(
    ["one\n", "two\n", "three\n", "four\n"],
    ["zero\n", "one\n", "tree\n", "four\n"],
    fromfile="Original",
    tofile="Current",
    )
    inspect(
    diff.join(""),
    content=(
    #|*** Original
    #|--- Current
    #|***************
    #|*** 1,4 ****
    #| one
    #|! two
    #|! three
    #| four
    #|--- 1,4 ----
    #|+ zero
    #| one
    #|! tree
    #| four
    #|
    ),
    )
    }

    diff_bytes

    fn diff_bytes(format : DiffFormat, a : Array[Bytes], b : Array[Bytes], fromfile? : Bytes, tofile? : Bytes, fromfiledate? : Bytes, tofiledate? : Bytes, n? : Int, lineterm? : Bytes, autojunk? : Bool) -> Iter[Bytes]

    Compares a and b, two sequences of lines represented as bytes rather than strings, producing a unified or context diff (see format) whose lines are also bytes.

    Inputs are losslessly converted to strings (one character per byte) so that files with unknown or inconsistent encodings can be compared, and the output is encoded back to bytes. This mirrors Python's diff_bytes; autojunk is forwarded to the underlying diff function.

    Example

    test {
    let diff = @difflib.diff_bytes(
    Unified,
    [b"\xa3odz is a city in Poland."],
    [b"\xc5\x81odz is a city in Poland."],
    fromfile=b"\xb3odz.txt",
    tofile=b"\xc5\x82odz.txt",
    lineterm=b"",
    )
    let expected : Array[Bytes] = [
    b"--- \xb3odz.txt", b"+++ \xc5\x82odz.txt", b"@@ -1 +1 @@", b"-\xa3odz is a city in Poland.",
    b"+\xc5\x81odz is a city in Poland.",
    ]
    assert_eq(diff.to_array(), expected)
    }

    get_close_matches

    fn get_close_matches(word : String, possibilities : Array[String], n? : Int, cutoff? : Double, autojunk? : Bool) -> Array[String] raise DiffError

    Uses [SequenceMatcher] to return a list of the best "good enough" matches for word among possibilities.

    • n (default 3): the maximum number of close matches to return; must be positive.
    • cutoff (default 0.6): a float in [0, 1]; possibilities that don't score at least that similar to word are ignored.
    • autojunk: the automatic junk heuristic of [SequenceMatcher].

    The best (no more than n) matches are returned sorted by similarity score, most similar first; ties are ordered like Python's heapq.nlargest on (score, string) pairs.

    Raises ValueError for an invalid n or cutoff.

    Example

    test {
    debug_inspect(
    @difflib.get_close_matches("appel", ["ape", "apple", "peach", "puppy"]),
    content=(
    #|["apple", "ape"]
    ),
    )
    let keywords = [
    "False", "None", "True", "and", "as", "assert", "async", "await", "break", "class",
    "continue", "def", "del", "elif", "else", "except", "finally", "for", "from",
    "global", "if", "import", "in", "is", "lambda", "nonlocal", "not", "or", "pass",
    "raise", "return", "try", "while", "with", "yield",
    ]
    debug_inspect(
    @difflib.get_close_matches("wheel", keywords),
    content=(
    #|["while"]
    ),
    )
    debug_inspect(@difflib.get_close_matches("Apple", keywords), content="[]")
    debug_inspect(
    @difflib.get_close_matches("accept", keywords),
    content=(
    #|["except"]
    ),
    )
    }

    is_character_junk

    fn is_character_junk(ch : Char) -> Bool

    Returns true for an ignorable character: a space or a tab.

    Example

    test {
    inspect(@difflib.is_character_junk(' '), content="true")
    inspect(@difflib.is_character_junk('\t'), content="true")
    inspect(@difflib.is_character_junk('\n'), content="false")
    inspect(@difflib.is_character_junk('x'), content="false")
    }

    is_line_junk

    fn is_line_junk(line : String) -> Bool

    Returns true for an ignorable line: one that is blank or contains a single '#' (after stripping whitespace).

    Example

    test {
    inspect(@difflib.is_line_junk("\n"), content="true")
    inspect(@difflib.is_line_junk(" # \n"), content="true")
    inspect(@difflib.is_line_junk("hello\n"), content="false")
    }

    ndiff

    fn ndiff(a : Array[String], b : Array[String], linejunk? : (String) -> Bool, charjunk? : (Char) -> Bool, autojunk? : Bool) -> Iter[String]

    Compares a and b (lists of lines) and returns a lazily generated [Differ]-style delta.

    charjunk defaults to [is_character_junk]; linejunk defaults to none.

    Example

    test {
    let diff = @difflib.ndiff(["one\n", "two\n", "three\n"], [
    "ore\n", "tree\n", "emu\n",
    ])
    inspect(
    diff.join(""),
    content=(
    #|- one
    #|? ^
    #|+ ore
    #|? ^
    #|- two
    #|- three
    #|? -
    #|+ tree
    #|+ emu
    #|
    ),
    )
    }

    restore

    fn restore(delta : Iter[String], which : Int) -> Iter[String] raise DiffError

    Extracts one of the two sequences that generated an [ndiff] / [Differ] delta: lines originating from sequence which (1 or 2), with the line prefixes stripped.

    Raises ValueError if which is not 1 or 2. Unlike Python, where the check happens when the generator is first advanced, it is raised immediately. The lines are produced lazily.

    Example

    test {
    // keep the delta in an array: an `Iter` can be consumed only once
    let diff = @difflib.ndiff(["one\n", "two\n", "three\n"], [
    "ore\n", "tree\n", "emu\n",
    ]).to_array()
    inspect(
    @difflib.restore(diff.iter(), 1).join(""),
    content="one\ntwo\nthree\n",
    )
    inspect(@difflib.restore(diff.iter(), 2).join(""), content="ore\ntree\nemu\n")
    }

    unified_diff

    fn unified_diff(a : Array[String], b : Array[String], fromfile? : String, tofile? : String, fromfiledate? : String, tofiledate? : String, n? : Int, lineterm? : String, autojunk? : Bool, color? : Bool) -> Iter[String]

    Compares two sequences of lines and returns the delta as a unified diff.

    The lines are generated lazily, like Python's generator: nothing is computed until the first line is requested, and the returned iterator can be consumed only once.

    • n (default 3): the number of context lines; must not be negative (it panics otherwise).
    • lineterm (default "\n"): appended to the control lines (---, +++, @@). Set it to "" for inputs without trailing newlines.
    • fromfile, tofile, fromfiledate, tofiledate: header fields.
    • color: emit ANSI colors similar to git diff --color. Unlike Python, the environment (NO_COLOR, tty detection, ...) is not consulted.
    • autojunk: the automatic junk heuristic of [SequenceMatcher].

    Example

    test {
    let diff = @difflib.unified_diff(
    ["one", "two", "three", "four"],
    ["zero", "one", "tree", "four"],
    fromfile="Original",
    tofile="Current",
    fromfiledate="2005-01-26 23:30:50",
    tofiledate="2010-04-02 10:20:52",
    lineterm="",
    )
    inspect(
    diff.join("\n"),
    content=(
    #|--- Original 2005-01-26 23:30:50
    #|+++ Current 2010-04-02 10:20:52
    #|@@ -1,4 +1,4 @@
    #|+zero
    #| one
    #|-two
    #|-three
    #|+tree
    #| four
    ),
    )
    }