<fragment><nonXMLChar><charCode>0</charCode></nonXMLChar></fragment>
<fragment><stringData>456</stringData></fragment>
<fragment><nonXMLChar><charCode>1</charCode></nonXMLChar></fragment>
<fragment><stringData>789</stringData></fragment>
<fragment><nonXMLChar><charCode>2</charCode></nonXMLChar></fragment>
<fragment><stringData>123</stringData></fragment>
</nonXMLString>

So our nonXMLString is of a type which is array of fragment, a fragment is a choice of either (legal XML) stringData, or a nonXMLChar.

The nonXMLChar has a child element because it will need to convert to from a string so will use inputValueCalc and outputValueCalc to do so, so it needs to be a sequence so that it can have the other hidden elements needed to pull this off.

stringData would have lengthKind="pattern" and a pattern that allows any sequence of XML-allowed characters.

nonXMLChar would have a hidden first child element of type string of explicit length 1 with an assertion that the string match a pattern that is any of the illegal characters (but just one of them). The charCode child element would inputValueCalc to get the character code of the character. For 8 bit encodings it would be ok as a table lookup in XPath, but for unicode..... we'd need a function that returns a character code.

If you just have one embedded illegal character, like NUL, then you could just model it as a separator, which would simplify things considerably (and is possible in a someday XML 1.1 future since NUL is then the only disallowed character.)

But for XML 1.0's illegal characters, we need to be able to convert to/from some non-string representation if we are to preserve information content. Hence we need these additional functions.