Repository navigation
Custom ext type registration #68
SeanTAllen
started this conversation in
Research
Replies: 1 comment
|
if we implement this, we should remember to add an example of its usage to examples/ |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
The problem
When working with MessagePack extension types, callers must manually manage the mapping between ext type codes and their serialization logic. The current API provides the raw building blocks --
MessagePackEncoder.extwrites an ext type code and byte payload,MessagePackDecoder.extreturns a(U8, Array[U8] val)tuple -- but the caller is responsible for dispatching on type codes in both directions.For encoding, the caller must remember which ext type code maps to which application type, serialize the application value to bytes, and call
extwith the correct code. For decoding, the caller reads the(code, data)tuple and writes amatchorif/elseifchain to dispatch to the correct deserializer:This works, but it has problems:
1andPointor2andColorexists only in the programmer's head and in scatteredmatcharms.A registration mechanism would centralize these mappings: define the (code, encoder, decoder) tuple once, then use it everywhere.
How other libraries handle this
Python (msgpack-python)
Python uses callback parameters on the packer and unpacker. The
defaultparameter onpackbhandles encoding, andext_hookonunpackbhandles decoding:Encoding and decoding hooks are separate parameters, not a unified registry. The caller is responsible for keeping them in sync. Unrecognized codes in
ext_hookfall through by returningExtType(code, data).Ruby (msgpack-ruby)
Ruby provides a
MessagePack::Factorythat centralizes registration:Registration is on the factory, not the individual packer/unpacker. The factory creates packers and unpackers that share the same type registry. The
packer:andunpacker:options default toto_msgpack_extandfrom_msgpack_extmethods on the registered class, but can be overridden with procs. Arecursive: trueoption passes the packer/unpacker to the callback for encoding nested msgpack structures.JavaScript (@msgpack/msgpack)
JavaScript uses an
ExtensionCodecclass with aregistermethod:The codec is passed as an option to
encode/decode. Theencodecallback returnsnullfor non-matching objects, which tells the encoder to try the next registered type. The codec is passed recursively when encoding nested structures.Go (vmihailenco/msgpack)
Go provides package-level registration functions:
The simplest path is
RegisterExt, which takes an ext ID and a value implementing bothMarshalerandUnmarshaler. Thevalueparameter is used for its type only (typically passed as a nil pointer:(*MyType)(nil)). Registration is global and expected to happen at init time. The library panics if the type-to-ID mapping is not bijective (no two types sharing an ID, no type registered under two IDs).C# (MessagePack-CSharp)
C# uses a formatter/resolver pattern. Custom types implement
IMessagePackFormatter<T>, and formatters are registered through a resolver chain:Formatters are composed into resolvers using
CompositeResolver. The resolver chain is set onMessagePackSerializerOptions, which is passed to serialize/deserialize calls. This is a general-purpose serialization framework, not specific to ext types.C++ (msgpack-cxx)
C++ uses compile-time adaptor specialization. The intrusive approach uses a macro:
The non-intrusive approach specializes template structs in the
msgpack::adaptornamespace. This is entirely compile-time -- there is no runtime registry. Custom types are serialized as msgpack arrays or maps of their fields, not as ext types.Common patterns
Across these libraries, several patterns emerge:
Registration is centralized. A single registry (factory, codec, or package-level function) holds all type-code mappings. Individual encode/decode calls reference the registry rather than containing dispatch logic.
Encode and decode are paired. Registration takes both an encoder and decoder for the same type code, ensuring they stay in sync.
The registry is configured at startup, then used immutably. Ruby's factory creates packers/unpackers from it. JavaScript passes the codec as an option. Go registers at init time and panics on conflicts. None of these libraries support modifying the registry while encoding/decoding is in progress.
Type erasure is unavoidable. The registry maps type codes to generic byte blobs. Decoding returns a dynamically-typed value that the caller must cast or match. Python returns
object, JavaScript returnsunknown, Go returnsinterface{}. The type safety boundary is at registration time, not at decode time.Design
The type erasure problem in Pony
The fundamental tension in ext type registration is between type safety and genericity. The registry must store handlers for different types under a uniform interface. In dynamically typed languages (Python, Ruby, JavaScript), this is natural -- everything is an object. In Go,
interface{}serves as the universal type. In Pony, the equivalent isAny val.A decode handler takes
Array[U8] val(the ext payload) and returns something. That something must be a type that the registry can store uniformly. In Pony, this means the handler's return type must be part of a union or must beAny val:Option B is impractical -- the union would need to include every possible application type, which defeats the purpose of extensibility. Option A works but forces the caller to
matchor useason the decoded value:This is the same pattern used by every other library. The type safety boundary moves from compile time to runtime at the point where the decoder returns a value. This is inherent to the problem: the registry doesn't know at compile time which type code will appear in the next byte of the stream.
The
Any valcapability constraintUsing
Any valas the type erasure boundary imposes a capability constraint: all values flowing through the registry must beval. This is actually appropriate for a serialization library:val.valsatisfies this.valregistry can be shared across actors, which matters for concurrent decoding.The constraint rules out encoding
refvalues directly. The caller would need to create avalsnapshot before encoding. This is consistent with how msgpack works generally -- you are serializing data, not mutable state.Interface design
The
encodemethod takesAny valbecause the registry dispatches based on ext type code, not on the Pony type of the value. The caller is responsible for passing the correct value to the correct handler. A handler implementation wouldmatchoras-cast the input:Handlers are
val(primitives are naturallyval, and class instances can be constructed asval). This means they can be stored in an immutable map and shared freely.Result types
The registry methods return union types instead of using bare
?, following the pattern established byMessagePackStreamingDecoder'sDecodeResult. This lets callers distinguish between "no handler registered" and "handler failed" without relying on context to diagnose errors.On the encode side, the result type is straightforward:
On the decode side, the success value is
Any val(type-erased). Naively writing(Any val | UnregisteredExtType | ExtDecodeError)collapses to justAny valbecause everyvaltype is a subtype ofAny val. A wrapper class resolves this:The caller unwraps
d.valueand then matches on the application type -- the same runtime cast that would be needed regardless, but now with explicit error discrimination at the registry layer.Registry design
The registry is a
valclass holding an immutable map from type codes to handlers:Builder pattern for construction
The registry is immutable after construction. A mutable builder accumulates registrations, then produces the frozen registry:
The builder errors on duplicate codes at registration time, following Go's approach of treating code conflicts as programmer errors. This catches a common mistake early rather than producing surprising behavior at encode/decode time.
The builder also validates that registered codes fall in the user-defined range (0-127). The MessagePack spec reserves codes -1 through -128 (unsigned 0xFF through 0x80) for predefined types -- timestamps already use ext type -1 (0xFF). The
MessagePackStreamingDecoderalready treats code 0xFF specially, returningMessagePackTimestampinstead ofMessagePackExt, so a user handler registered for 0xFF would silently never be reached through the streaming decoder path. Enforcing the range prevents this class of bugs. Callers who need custom handling for reserved codes can use the low-levelMessagePackEncoderandMessagePackDecoderprimitives directly.Interaction with existing APIs
The registry is a companion to the existing primitives, not a replacement. The stateless
MessagePackEncoderandMessagePackDecoderremain unchanged. The registry provides convenience on top of them.With
MessagePackDecoder: The caller decodes an ext value usingMessagePackDecoder.ext(reader)?to get the(code, data)tuple, then passes it to the registry:With
MessagePackStreamingDecoder: The streaming decoder already returnsMessagePackExtfor ext values, which containsext_typeanddatafields. The caller passes these to the registry:In both cases, the registry is used after the low-level decoder has done its work. The registry does not need to be integrated into the decoder itself -- it operates on the already-decoded
(code, data)pair.For encoding: The registry wraps the encoder to handle the ext header and payload assembly:
Encode-side asymmetry: The registry fully solves decode-side dispatch -- the ext type code arrives from the wire and the registry maps it to the correct handler automatically. On the encode side, the improvement is the centralization of serialization logic in the handler, but the caller must still know which ext code maps to which application type (e.g., that
Pointuses code 1). This asymmetry is inherent: decode dispatch is data-driven (the code is in the wire format), while encode dispatch is caller-driven (the application knows its own types). This matches how all surveyed libraries work -- none of them provide fully automatic encode-side dispatch without the caller specifying which type to encode as.Encoding overhead: payload assembly
The
encode_extmethod uses a temporaryWriterfor the handler to write payload bytes into, then collects the writer's chunks into a singleArray[U8]before passing it toMessagePackEncoder.ext. This involves an extra allocation and copy becauseextneeds the total payload size upfront to write the ext header.For small payloads (the common case for ext types -- UUIDs, coordinates, custom scalars), this overhead is negligible. For large payloads, the copy is proportional to the payload size. An alternative design could have the handler report its payload size before encoding, allowing the registry to write the ext header first and then let the handler write directly to the output writer. This would avoid the intermediate buffer but complicate the handler interface. The simpler approach (temporary writer + collect) is proposed as the starting point, with the optimized path available as a future enhancement if profiling shows the overhead matters.
Why the registry is separate from the decoder
The registry could be integrated into the streaming decoder, so that
next()returns application types directly instead ofMessagePackExt. This was considered and rejected for several reasons:It would change
DecodeResult. The streaming decoder's return type is(MessagePackValue | NotEnoughData | InvalidData). AddingAny valto this union would make everymatchon the result less precise -- the caller would need to handle an opaque type alongside the well-typed msgpack primitives.It couples the decoder to application types. The streaming decoder is a general-purpose msgpack decoder. Coupling it to a specific set of application-defined handlers makes it less reusable.
The composition is trivial. Using the registry as a post-processing step on the decoder's output is straightforward and doesn't sacrifice ergonomics. The two-step pattern (decode ext, then dispatch via registry) is explicit about what happens at each stage.
Thread safety and capability model
The registry is
valafter construction. This means:Encoding requires a
Writer refparameter, which is inherently single-actor. Each actor that encodes needs its ownWriter. This is the same constraint as usingMessagePackEncoderdirectly -- the registry does not change it. The sharing benefit is for the registry itself and for decode operations, which create new values from data.The builder is
refduring construction, ensuring that registration happens in a single-actor context where mutation is safe. Thebuild()method produces an immutable registry from the builder's mutable state, using arecoverblock to lift the handler map fromreftoisoand then toval. The builder clears its internal state after building, so a second call tobuild()produces an empty registry. This is the standard Pony pattern for constructing shareable data structures.Alternative: per-type decode without
Any valAn alternative design avoids
Any valby providing typed decode methods on the registry:This doesn't actually help. The registry stores handlers as
ExtTypeHandlervalues, so the decode method returnsAny valinternally. The generic method would just unwrapExtDecodedand add anascast:The caller can do this unwrap and cast themselves. Adding a generic wrapper provides no additional type safety -- the
ascast can still fail at runtime -- and adds API surface without value.Scope
This proposal covers:
ExtTypeHandlerinterface for implementing ext type encode/decode logicExtTypeRegistryclass for storing handler mappingsExtTypeRegistryBuilderclass for constructing registriesThis proposal does not cover:
MessagePackEncoder,MessagePackDecoder, orMessagePackStreamingDecoderto accept registries directlyAll reactions