Skip to content

Add an optional GLiNER classifier sidecar for custom routing #735

Description

@pst2154

Summary

Add and production-harden an optional GLiNER classifier sidecar for Switchyard's existing custom multi-target router. The sidecar should classify requests into deterministic-tool, small-model, reasoning-model, or human-review groups while Switchyard continues to own schema validation, fallback ordering, session affinity, and final dispatch.

Why

A deployment-local classifier can reduce routing latency and support offline or domain-tuned operation without adding Python or PyTorch to the Switchyard server. It should remain optional and communicate through the existing OpenAI-compatible classifier boundary.

Prototype evidence

A zero-shot prototype using the public fastino/gliner2.5-base-v1 checkpoint was evaluated on 32 scored routing cases plus four ambiguous inspection cases:

Classifier Overall accuracy Ordinary cases Adversarial cases Median latency
GLiNER, single A100 80 GB 78.1% 87.5% 50.0% 20.1 ms
GLiNER, local CPU 78.1% 87.5% 50.0% 54.6 ms
TypeSafe Jev, managed remote call 96.9% 100% 87.5% 266.3 ms

The latency rows include different deployment topologies and are not hardware-normalized model-execution comparisons. The accuracy suite is small and synthetic. Several incorrect GLiNER predictions had high confidence, so a global confidence threshold reduced accuracy rather than improving it.

The complete Switchyard path was also exercised against mock completion targets: classifier request, structured verdict, policy selection, and downstream dispatch all worked for the four route groups.

Proposed shape

  • Reuse llm_classifier with mode = "custom".
  • Keep GLiNER in a separately deployed OpenAI-compatible sidecar.
  • Use internal semantic labels that differ from public target names, then map them to Switchyard groups.
  • Validate the target enum supplied by Switchyard before returning a verdict.
  • Do not log prompt content or credentials.
  • Keep consequential actions behind deterministic checks or human approval regardless of classifier confidence.

Production-hardening work

  • Expand evaluation with representative, domain-specific traffic and cost-weighted confusion metrics.
  • Fine-tune or otherwise adapt GLiNER for the intended routing taxonomy.
  • Calibrate uncertainty and fallback policy per route rather than using one global threshold.
  • Add batching, concurrency, saturation-throughput, and soak benchmarks.
  • Define deployment health, readiness, observability, and model-version pinning.
  • Add prompt-injection regression cases and verify that route labels cannot authorize actions.
  • Document the optional sidecar configuration and reproducible benchmark procedure.

Non-goals

  • Adding GLiNER/PyTorch as mandatory Switchyard dependencies.
  • Treating model confidence as permission for destructive or high-stakes actions.
  • Claiming production readiness from the current synthetic benchmark.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions