Diff and Patch#

Diffing and Patching YAML

WIP

Goals#

YTool diffing support strives to

  1. Create succint but comprehensive and readable yaml diffs a. include LCS computation on arrays b. include string-to-string diffs
  2. Work with yaml tags
  3. Interoperate with the existing patching mechanism a. work for !key(name) -- see Keyed Arrays
  4. Include string-to-string diffs

Format#

Basic Format#

a:
  b: 
  # this indicates there's an array in from and to
  # here which differs

  !arraydiff  

  # differs one at index 172 (hex 0xac)
  # with a replacement
  00ac: !replace
    from: w
    to: x

  # delete v in from 
  c: !delete v
  # insert x in to
  d: !insert x

  # line-by-line strdiff
  e: !strdiff(true)

  # char-by-char strdiff
  f: !strdiff(false)

Anywhere there is a diff we can generate a forward patch by eliding from and embedding to where the diff node is. Likewise in reverse, since a diff can be reversed.

What a key means#

The keys of an !arraydiff or a !strdiff are positions in the sequence the two sides share, not offsets into either one of them. Every unit of either side takes one position: an unchanged unit takes the same position in both, a deleted one is the from's alone, an inserted one the to's alone, and a !replace takes as many as its longer side.

The unit is the element for an !arraydiff, and for a !strdiff it is whichever the argument names -- a line under !strdiff(true), where the lines are those of splitting on \n, and a rune, never a byte, under !strdiff(false).

# alpha         ->  alpha        the from's beta, gamma and delta take
# beta              BETA         positions 1, 2 and 3, so what follows
# gamma             epsilon      them starts at 4 whichever side it is
# delta             ...          read from
# epsilon
# ...
!strdiff(true)
1: !replace
  from: |-
    beta
    gamma
    delta
  to: BETA

Reversing a diff swaps every from with its to and rewrites no key, so a position has to mean the same thing read in either direction. That is why a delete counts and why a !replace counts by its longer side: either one measured from the result alone would move under reversal and throw off every key after it.

Tag Format#

In the basic format, tags are preserved in any !replace operation,

Keyed Arrays#

When both sides of a diff carry !key(f) with the same field, elements are matched by their key rather than by position, and the diff names the elements which changed:

items: !key(sku)
- !insert
  sku: G
  q: 3
- !delete
  sku: W
  q: 1

Applying that to a list which has since gained or lost other elements leaves those alone, which is the point: position-based diffs cannot do this. A reordering of the same elements is not a difference at all.

Two properties worth stating, since neither is obvious:

  • The keyed branch is taken only when both sides carry !key(f) with the same field, the field is a resolvable path, and each element has it. Any miss falls back to a positional array diff, silently -- so two documents keyed by different fields diff by position rather than reporting a conflict.
  • Patch(a, Diff(a, b)) reproduces b's content, but not necessarily node-for-node: object fields are normalised to alphabetic order, and presentation tags such as !bracket are dropped when a value is created rather than merged into an existing one. Compare with a diff, not with a deep equality.

String Diffs#

String diffs are computed rune by rune unless both sides have multiple lines, in which case they are computed line by line.

When strings differ and the length of the text in the differences is at least half the minimum length of the source and target text, the texts are simply replaced wholesale w/out a string diff. This is a heuristic that will need improving over time, as some strings are really representing enums and the like under the hood and here in-string diffs don't help readability.

String diffs do not apply to fields.

Comments#

A diff is about values and says nothing about comments unless asked: Diff is blind to them, and two documents differing only in what was said about a value diff to nothing. DiffWith(a, b, DiffComments(true)) asks the other question -- are these the same document -- and answers with !comment, which states what the comments at a node are rather than replacing the value they describe (see matchpatch.md).

A patch answers with data: the result carries no comments, neither the patch's nor the ones the document being patched came in with, unless the patch is given mergeop.Comments(true). That is a deliberate policy rather than two accidents -- a head comment is a wrapper anything descending through discards, while a line comment rides on the node and every clone carries it, so "off" once meant keeping half of them.

The round trip holds either way: Patch(a, DiffWith(a, b, DiffComments(true)), Comments(true)) is b, comments included.

Matching is blind to comments in both directions and has no option: a match asks about the value and sees through what was said about it. The format says comments should be matchable and the tooling does not yet do it -- issue 8241kcggh12krgh4g1n0.

In o, all of this is -c: o diff -c, o patch -c, and o get -c / o list -c for answering with the node as it stands rather than the value a path names.