After some cleaning up and finding answers to remaining questions this will be useful to others writing ubjson parsers and generators, and I intend to get it to that state around the time of the next draft or sooner. Right now these are just my notes. Help looking for misunderstandings will be appreciated.
; UBJSON specification draft [???version???]
; copyright information / freedom to share / no warranty goes here
document = object ;
; [OPINION] If segments of UBJSON are stored inside another file format that has a way to detect damage, it’s acceptable to
; omit the initial "{", but it must be added in front when the UBJSON is extracted to a separate file or stream.
; -> [PARSER OPTION] use <document = objectcontent> instead.
item = object | array | string | number | singlebyte ;
; noop is not an item
noop = "N" ;
singlebyte = boolean | null ;
boolean = "T" | "F" ;
; null, also infinity
null = "Z" ;
number = "i", ivalue
| "U", Uvalue
| "I", Ivalue
| "l", lvalue
| "L", Lvalue
| "d", dvalue
| "D", Dvalue
| "H", Hvalue
;
string = "S", Svalue
| "C", Cvalue
;
; The comments on strong typing will explain the reason for the <*value> rules.
ivalue = ? int8-be ? ;
Uvalue = ? uint8-be ? ;
Ivalue = ? int16-be ? ;
lvalue = ? int32-be ? ;
Lvalue = ? int64-be ? ;
dvalue = ? IEEE 754 binary32 ? ; [INVESTIGATE] is be/le conversion needed?
Dvalue = ? IEEE 754 binary64 ? ; [INVESTIGATE] same
Cvalue = ? ASCII character ? ;
; <Svalue> and <Hvalue> are below <count>.
;
;
; Lengths
; =======
count = "i", ? int8-be: value >= 0 ?
| "U", ? uint8-be ?
| "I", ? int16-be: value >= 0 ?
| "l", ? int32-be: value >= 0 ?
| "L", ? int64-be: value >= 0 ?
;
; [OPINION] "i" should be accepted in <count> but "U" should be produced instead.
;
; 32 bit systems must not cast int64 down to int32 in malloc, that would cause buffer overflows.
; [OPINION] they may fail instead, with an out of memory error or similar.
Hvalue = count, ? decimal number represented by ASCII string of length int(count) matching
[ "-" ],
( "0" | ( digit - "0", { digit } ) ),
[ ".", digit, { digit } ],
[ ( "e" | "E" ), [ "+" | "-" ], digit, { digit } ]
? ;
; By definition int(count) must be > 0 here, it might be good idea to check that before trying to parse the high precision number.
; Parsers that cannot interpret high precision numbers do not have to verify them, they should fail when they encounter one
; unless:
; [PARSER OPTION] pass high precision numbers as-is
; [PARSER OPTION] skip high precision numbers (WILL NOT IMPLEMENT)
Svalue = count, ? UTF-8 string represented by int(count) bytes ? ;
;
;
; Strong Typing
; =============
;
; Containers can have a type, which is represented by one of the following characters. For each, there is a matching
; <$(character)value> rule that can be used in the container content (and which is reused in some other places).
type = "i" | "U" | "I" | "l" | "L" | "d" | "D" | "H" | "S" | "C" | "{" | "[" ;
; Typed containers cannot contain typed containers themselves, so we need a way to disable the type system. Parsers may of
; course use a different method than through the grammar like we do here.
; <simple$(container)>s contain <simpleitem>s instead of <item>s, and <simpleitem>s cannot be not-simple <$(container)>s.
; [QUESTION] Are additional "#" meant to be allowed in contained <object>s and <array>s, unlike additional "$" ?
; [value
arrayvalue = { simpleitem, { noop } } "]"
| "#", { noop }, count, { noop }, ?int(count)? * ( simpleitem, { noop } )
;
; {value
objectvalue = { Svalue, { noop }, simpleitem, { noop } }, "}"
| "#", { noop }, count, { noop }, ?int(count)? * ( Svalue, { noop }, simpleitem, { noop } )
;
simpleitem = simpleobject | simplearray | item - ( object | array ) ;
simplearray = "[", { noop }, arrayvalue ;
simpleobject = "{", { noop }, objectvalue ;
; Finally, the full featured containers.
; [OPINION] zero length containers should be produced without a <count> and <type>.
array = "[", { noop }, arraycontent ;
arraycontent = { item, { noop } }, "]"
| "#", { noop }, count, { noop }, ?int(count)? * ( item, { noop } )
| "$", { noop }, type, { noop }, "#", { noop }, count, { noop }, ?int(count)? * ?$(type)value?, { noop }
| "$", { noop }, singlebyte, { noop }, "#", { noop }, count, { noop }
;
object = "{", { noop }, objectcontent ;
objectcontent = { Svalue, { noop }, item, { noop } }, "}"
| "#", { noop }, count, { noop }, ?int(count)? * ( Svalue, { noop }, item, { noop } )
| "$", { noop }, type, { noop }, "#", { noop }, count, { noop }, ?int(count)? * ( Svalue, ?$(type)value?, { noop } )
| "$", { noop }, singlebyte, { noop }, "#", { noop }, count, { noop }, ?int(count)? * ( Svalue, { noop } )
;
; [OPINION] like in JSON keys must not be repeated, but ignoring that issue might be acceptable for stream parsers.
; -> [PARSER OPTION] no key reuse checking
While preparing for https://github.com/tbuitenhuis/zgio I’ve been writing an EBNF specification of ubjson. I’ve been asked to post it here.
After some cleaning up and finding answers to remaining questions this will be useful to others writing ubjson parsers and generators, and I intend to get it to that state around the time of the next draft or sooner. Right now these are just my notes. Help looking for misunderstandings will be appreciated.