Lance Table Schema Semantic Type Contract #7073
Replies: 2 comments 1 reply
|
This proposal is an extension from @wjones127's #5864 |
|
+1, this is a great writeup (though it is a lot of information 😆). Is this a proposal or just discussion? First, I agree with almost everything here. This would be a good system.
I think you are arguing that we need to centralize on one type before we send it to the writer. For example, if the user has a However, what happens if the table format picks On the other hand, if we pick Why can't we just give the writer whatever it is the user gave us? If they give us Or am I misunderstanding something here? |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Summary
This design makes
Field.logical_typein the table schema name a column's semantic type: the value domain and query semantics that every engine must preserve. Arrow layout choices that do not change values (32- or 64-bit offsets, view layouts, dictionary encoding, decimal width) leave the logical type. The writer encodes whatever valid representation it receives. The reader returns the layout named by an advisorylance-schema:output-encodingfield hint or by a request override. Only value-transforming types (json,blob) convert between input and storage. Applications can define their own semantic types as Arrow extension types on top of a core type. Readers that do not recognize an extension read it as its storage type, so new application types need no format vote.Engines get a type system they can map without knowing Arrow encodings. Users can append
LargeUtf8,Utf8View, or dictionary data to astringcolumn, and a read returns the layout they wrote with. Lance can add view types and pick storage layouts without format changes. The contract is table-level only: data files, their encodings, andLanceFileVersionare unchanged.Problem
logical_typecurrently maps one-to-one to an Arrow type, so a single string mixes three concerns:string,json,decimalprecision and scalestringvslarge_stringdict:string:int16:false,decimal:128:10:2vsdecimal:256:10:2This has concrete costs:
LargeUtf8data to astringcolumn fails even though every value fits.Proposed Design
Semantic Types
A semantic type is a distinct type exactly when it changes the value domain, precision, or computation semantics. Layout and encoding choices never create a new type. By this rule,
int32andint64remain distinct because width bounds the values.floatanddoubleremain distinct because width sets arithmetic precision.stringandlarge_stringare one type.Core semantic types fall into three classes.
Unchanged types. These keep their current
logical_typestring, and each has exactly one Arrow representation:null,bool,int8…int64,uint8…uint64,halffloat,float,double,date32:day,date64:ms,time32:*,time64:*,duration:*,timestamp:*,struct,map,fixed_size_list:<elem>:<n>(includinglance.bfloat16elements), andfixed_size_binary:<n>.Representation-only types. Several Arrow layouts hold the same values. The writer encodes any listed physical layout unchanged, and each data file may use a different one.
logical_type)stringUtf8,LargeUtf8,Dictionary<int*, Utf8 | LargeUtf8>Utf8Viewutf8,large_utf8,utf8_view,dictionary:<key>:<value>binaryBinary,LargeBinary,Dictionary<int*, Binary | LargeBinary>BinaryViewbinary,large_binary,binary_view,dictionary:<key>:<value>listList,LargeListlist,large_listdecimal:<p>:<s>Decimal128(p, s),Decimal256(p, s)decimal128/decimal256that holdsp, then the otherDictionary encoding is a layout of
stringorbinaryvalues, not a type. A view input is written as a non-view layout that holds every value; the writer chooses which one.Value-transforming types. Input values are converted to a different stored representation, and output converts them back. Only these types need conversion logic in the writer and reader.
jsonarrow.jsonoverUtf8,LargeUtf8, orUtf8View(validated, then converted to JSONB);lance.jsonoverLargeBinary(unchanged)LargeBinary+lance.json(JSONB)arrow.json(text),lance.json(JSONB)blobLargeBinary+lance-encoding:blob. Blob v2:struct+lance.blob.v2This design only classifies
blob. Its representations and read modes remain governed by the Blob specification.Output Encoding
lance-schema:output-encodingis an optional field metadata entry that names the Arrow layout a read returns for that field. Its values are the output encodings listed in the tables above. For nested types, each child field carries its own entry. A reader picks the output layout in this order:lance-schema:output-encodingentryWhen a writer creates a column (table creation, add column, or alter type), it records the input Arrow layout as the field's output encoding if that layout differs from the default. This way a
LargeUtf8input reads back asLargeUtf8. Appends do not change the entry. Changing the entry is a metadata-only schema update.The output encoding is advisory. It never affects which values are stored, schema compatibility, or correctness. A reader that does not recognize a value, or cannot produce that layout, uses the type default instead.
Extension Semantic Types
An extension semantic type is a core field that carries
ARROW:extension:nameand, optionally,ARROW:extension:metadata. The core type is its storage type. Example:logical_type = "fixed_size_list:float:4"withARROW:extension:name = "example.bbox". A field carries one extension name, so an extension cannot yet be layered on a type that already uses one (json, Blob v2,lance.bfloat16elements); see open decision 2.lance.are reserved for Lance. A Lance-owned extension may transform values only if a format version or feature flag gates it;lance.blob.v2requires data storage version 2.2. Other extensions should use a project prefix such aslancedb.or a reverse-DNS name.Schema Compatibility
Two fields are compatible for append, merge, and update when their semantic types and semantic parameters are equal. Output encodings, physical layouts, and dictionary keys are not compared. Consequences:
Decimal256(10, 2)todecimal:10:2succeeds.Decimal128(12, 2)todecimal:10:2fails, because the value domain differs. Changing precision is a type change, not a compatible append.Field IDs, nullability, and nested structure follow the existing rules.
Data Files and the Physical Schema
A data file's schema records the exact physical layout encoded in that file, using the existing Arrow-mapped strings (
large_string,dict:string:int16:false,decimal:128:10:2, …). These strings keep their current meaning inside data files. The table schema is the only source of semantics: readers resolve semantics through field IDs in the table schema and never infer them from a data file's physical layout. The encodings, footer, page metadata, and global buffers of data files are unchanged.Contract and Invariants
struct,list, andmapchildren.Failure Semantics
jsoninput is not valid JSONutf8value over the 32-bit offset limit)large_utf8orutf8_view. Readers may emit smaller batches to fit, but must never truncate.lance-schema:output-encodingto a value that is invalid for the field's type or parameters (for example,decimal128for precision 40)lance-schema:output-encodingvalueCompatibility and Migration
Feature flag. A table that follows this contract sets a new feature flag,
FLAG_SEMANTIC_TYPES, in bothreader_feature_flagsandwriter_feature_flags. The spec PR assigns its bit. Once set, the flag remains set in all later versions, including restores.Without the flag, older clients would misread these tables in ways that are hard to detect:
stringas exactlyUtf8, ignore the output encoding, and return different Arrow types than the table specifies.decimal:<p>:<s>at all (evidence: the currentTryFrom<&LogicalType>requires four:-separated parts).The flag turns all of these into one clear "unsupported" error.
Legacy tables (flag not set) keep their current behavior exactly. Each
logical_typenames one Arrow type, schema compatibility compares those types, and no output encoding applies. New readers do not rewrite legacy schemas.Enablement. New tables created with data storage version 2.3 or later set the flag (open decision 1). Tables on earlier storage versions stay legacy unless explicitly upgraded.
Upgrade. Data files do not change, so upgrading an existing table is a single metadata-only commit. The commit rewrites each legacy alias to its canonical type plus the output encoding that reproduces the current read type, then sets the flag. After the upgrade, reads return the same Arrow types as before. There is no downgrade.
logical_typestring,binary,listlist.structlist(the child field is alreadystruct)large_stringstringlarge_utf8large_binarybinarylarge_binarylarge_list,large_list.structlistlarge_listdict:<value>:<key>:false<value>dictionary:<key>:<value layout>decimal:128:<p>:<s>decimal:<p>:<s>p <= 38, so the default is alreadydecimal128)decimal:256:<p>:<s>decimal:<p>:<s>decimal256whenp <= 38; none otherwiseWriters of flagged tables write only canonical names in the table schema. Readers of flagged tables still interpret a legacy alias as its canonical type plus the implied output encoding.
Validation
This contract is accepted when all of the following are observable:
Utf8,LargeUtf8,Utf8View, andDictionarydata to onestringcolumn succeed across fragments. Reads return the recorded output encoding, and request overrides change it.json(struct<json>,list<struct<json>>,map<string, json>) writes JSONB and reads back asarrow.jsonthrough create, append, update, merge-insert, and take.Open Decisions
stable_file_version()is 2.2), and 2.3 data files already require new readers, so default-on adds no new compatibility break. Explicit opt-in only would leave most new tables on the legacy contract. Closing this needs PMC agreement.ARROW:extension:name = "lance.blob.v2". Recommendation: record extension semantic types under a Lance-owned key,lance-schema:extension-name(pluslance-schema:extension-metadata). On output, the reader maps these keys toARROW:extension:nameandARROW:extension:metadataonly when the field has no Arrow extension of its own. The alternative, reusing only the Arrow slot, cannot express media types over Blob v2. Closing this needs agreement with #8739 (structured Blob v2 descriptor extensions) on where media metadata lives.Implementation Notes
The spec change touches
docs/src/format/table/schema.mdandversioning.md. The implementation keeps semantic-type dispatch internal to the Rust core, with no public plugin trait. The write path gets a single boundary check, extending today'sSchemaAdapter. Reads choose layouts through the existingReaderProjection, which is designed to decode a column into another in-memory layout of the same semantic type when the decoder supports it; conversions a decoder lacks (for example, between decimal widths) run after decoding. The writer boundary check changes no format and can land before the spec vote. The upgrade commit is a small metadata rewrite and is planned for the first implementation, so existing tables can adopt view types and relaxed appends.All reactions