Skip to content

Support typed heterogeneous result sets from one YQL request #38

Description

@asmyasnikov

Applications often need several independent read results for one response: a count, a page of records, and a status flag. Sending these SELECT statements as one authored YQL request can avoid a separate execution round trip for each result. Generate one method that submits the complete request and returns an aggregate with separately typed result sets in statement order, even when their column schemas differ. Every output is a typed collection of rows.

Proposed annotation

Introduce one opt-in query annotation, -- name: QueryName :multi. Infer each result set's column names and types from its SQL statement:

-- name: GetSummary :multi
SELECT 1 AS value;
SELECT "2"u AS value;
SELECT false AS value;

The scalar example deliberately produces three different schemas. Real applications can replace these expressions with a count query, a bounded collection query, and a Boolean status query. Explicit column aliases retain the existing contract for computed projections.

Generate stable ordinal fields for logical result sets. Every field is a typed slice/list, including results from scalar SELECT expressions and queries containing LIMIT 1. An illustrative Go API is:

func (q *Queries) GetSummary(ctx context.Context) (GetSummaryResult, error)

type GetSummaryResult struct {
    Result1 []GetSummaryResult1Row
    Result2 []GetSummaryResult2Row
    Result3 []GetSummaryResult3Row
}

Result1, Result2, and Result3 follow logical result-set order. Their row types have the columns and YQL types inferred by the analyzer; there is no untyped catch-all result array.

This is a proposed API, not an implemented annotation. Existing -- name: ... :one, :many, and :exec commands and their generated APIs keep their current behavior.

Initial scope and semantics

  • Start with a statically known sequence of two or more supported SELECT statements, shared DECLARE parameters, and a fully buffered typed aggregate. Resolve each statement in the shared analyzer; retain direct parser contexts and one resolved compilation result for all generators. Do not add generator-specific SQL analysis.
  • Send the entire authored script through one logical execute call per attempt. Never execute the SELECT statements separately, concatenate client-side results from separate calls, or retry an individual statement. Preserve the chosen runtime's transaction ownership, consistency mode, and caller-controlled retry policy; a single request alone must not be documented as guaranteeing a particular snapshot or atomicity level.
  • Derive the ordered output schema from top-level result-producing statements in the shared analyzer. Column aliases provide field names within each row; validate generated names and collisions using each target's normal rules. Nested SELECTs are not independent outputs. Ordinals count logical result sets, not DECLARE statements, other non-result-producing statements, or transport chunks.
  • Always return every result set as a typed collection containing all of its rows. Zero, one, or many rows are valid for every output; scalar expressions and LIMIT 1 do not change the generated collection type or introduce a row-count check. Preserve an empty logical set in its ordinal field so later sets never shift. Apply relevant target collection options, such as Go's empty-slice policy, consistently.
  • Match logical result sets to the inferred outputs by SDK result-set index/order using the actual pinned SDK/driver contract. A result set is not a transport message, response part, or row chunk: multiple chunks for one statement must be combined into the same output collection. Verify how each adapter represents empty sets and result-set boundaries before claiming support.
  • Missing or extra logical result sets, incompatible result schemas, decoding failures, server errors, and cancellation are errors. A valid empty set is not a missing set. Never report a partially populated aggregate as a successful result. Consume all result sets and validate the terminal server status even after all expected rows have been read; include the result-set index in errors when available.
  • Close or cancel every owned result/cursor on success and on every failure path, including a failure after earlier result sets were decoded. Keep caller-owned connections and transactions open. Unsupported runtime profiles must fail generation with an actionable diagnostic rather than discard all but the first set.

Delivery and validation

  1. Specify the single query annotation, ordinal aggregate naming, inferred row schemas, parameter scope, and unsupported-profile diagnostics; add focused analyzer and generation tests.
  2. Implement native Go and database/sql against their pinned SDKs, verifying real logical-result iteration APIs. Fully materialize every output collection before returning. Document that memory use grows with the total returned data; incremental APIs are a separate feature.
  3. Execute generated clients against pinned local-ydb. Cover the three distinct scalar schemas above; collection results with zero, one, and many rows; nullable columns; empty first/middle/last sets; a large set delivered across multiple server response parts; scalar and LIMIT 1 queries that still produce collection fields; correct advancement after every logical result set; cancellation and decoding failure after an earlier set; late server failure; and owned-resource cleanup. Add focused protocol/driver harness tests for missing, extra, or mismatched sets when valid static SQL cannot naturally produce them. Verify that all statements are submitted together, without per-statement execution calls.
  4. Document supported profiles and exact transaction/retry behavior, then extend other generators only after verifying their runtime APIs. Preserve the existing single-result acceptance suite.

Relationship to other work

Issue #33 covers scripts with multiple DML statements and at most one typed result set, and explicitly excludes heterogeneous multiple result sets. This issue covers several separately typed outputs from one request. The initial read-only implementation need not wait for mixed DML scripts; combining these capabilities can follow once both contracts are established.

Issue #35 covers incremental row consumption and resource ownership. This issue's first version buffers its outputs and does not depend on a streaming public API. Internal SDK transport streaming must still be handled correctly.

Keep dynamic/conditional result counts, incremental public iterators, arbitrary control flow, and mixed DML/RETURNING scripts outside this initial feature. This issue records future work; no implementation is requested by creating it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions