Skip to content

cloud_service: make DP report async on the mqtt, lan and ble channels - #723

Open
jianning773 wants to merge 4 commits into
tuya:masterfrom
jianning773:feature/dp-async-report
Open

jianning773 wants to merge 4 commits into
tuya:masterfrom
jianning773:feature/dp-async-report

Conversation

@jianning773

Copy link
Copy Markdown
Contributor

Problem

DP reports were sent inline on the caller's thread: every report blocked on the transport write — the mqtt TLS publish, the per-session LAN socket writes, and the BLE GATT transfer (subpackets with 20ms sleeps in between). Downlink DP commands are parsed and echoed on the high-priority work queue (wq_highpri), so one slow client or congested link stalled the whole business pipeline behind it, and the inline mqtt publish also raced with MQTT_ProcessLoop on the shared client context.

Change

One pattern applied to all three channels: the caller only validates, deep-copies and enqueues; the channel's own context drains and sends.

  • mqtt (adbafd2): hand the obj/raw reports to the existing async publish API; the mqtt loop task publishes and notifies completion via callback.
  • lan (e0b0d91): tuya_lan_dp_report_async/_flush backed by a 16-slot queue; the packed json string is copied on enqueue and drained by the lan_sock_loop thread each loop. The sock-loop select timeout is shortened to 100ms so a queued send leaves within 100ms.
  • ble (7589742): tuya_ble_dp_report_async/_flush backed by a 16-slot queue; the queued copy owns the dps array and any PROP_STR/raw payloads (deep-copied on enqueue, freed after send). Drained on WORKQ_SYSTEM — BLE has no thread of its own.

A full queue drops the report with a warning; dp sync covers the loss, matching the mqtt channel's existing timeout-drop semantics.

Verification (T5AI / BK7258 chat-bot board)

  • Build passes; the four new symbols are present in the flashed ELF.
  • LAN: 18/18 reports show the send moved off the caller — log order flipped from lan channel report → LAN send → mqtt channel report to lan channel report → mqtt channel report → LAN send; zero queue-full drops, zero socket faults, session lifecycle (30s heartbeat close) unchanged.
  • MQTT (regression): 20 PUBACKs covering all reports, no loss, behavior unchanged.
  • BLE: verified by build + elf symbols. The runtime path requires a cloud-offline BLE-direct session (the sdk stops BLE advertising while the cloud link is up), so end-to-end BLE testing is left as future work; cloud-online scenarios show no regression.
  • 5-minute soak with panel volume operations (19 downlinks + echoes): no crashes, no errors, heap and watchdog normal.

🤖 Generated with Claude Code

jianning773 and others added 4 commits September 23, 2026 14:33
Switch tuya_iot_dp_obj_report from tuya_iot_dp_report_json_with_notify
(inline QOS1 publish on the caller thread) to the _async variant, so the
TLS/TCP write is done by the mqtt loop task instead of blocking the
caller (typically WORKQ_HIGHTPRI handling the DP-receive echo) and no
longer races with MQTT_ProcessLoop on the shared mqtt client context.

The completion path is unchanged: PUBACK -> dp_sync_cb marks the DPs
PV_STAT_CLOUD, timeout/failure schedules the PV_STAT_LOCAL resync.

Also give publish_list a real mutex (the LOCK/UNLOCK markers were
placeholders with no actual lock): producers append from caller threads
while the mqtt loop task flushes and completes entries. Completion
callbacks are invoked outside the critical section so a callback that
publishes again cannot deadlock.

Co-Authored-By: Claude Code <noreply@anthropic.com>
LAN DP reports were sent inline on the caller's thread, blocking
wq_highpri for the per-session socket writes. Queue the packed json
strings on a dedicated queue and let the lan_sock_loop thread flush
them, so the caller only pays for the enqueue.

- tuya_lan.c: add tuya_lan_dp_report_async/_flush backed by a 16-slot
  queue; create it in tuya_lan_init, drain and release it in
  tuya_lan_exit
- lan_sock.c: flush the queue in tuya_sock_loop_run and shorten the
  select timeout to 100ms so a queued send leaves within 100ms
- tuya_iot_dp.c: hand the lan obj/raw reports to the async queue

A full queue drops the report with a warning; dp sync covers the loss,
matching the mqtt channel's timeout-drop semantics.

Co-Authored-By: Claude Code <noreply@anthropic.com>
BLE DP reports ran inline on the caller's thread and blocked it for
the whole GATT transfer (20ms sleeps between subpackets). Queue a
deep copy of the report and let WORKQ_SYSTEM flush it; BLE has no
thread of its own.

- ble_dp.c: add tuya_ble_dp_report_async/_flush backed by a 16-slot
  queue; the queued copy owns the dps array and any PROP_STR/raw
  payloads (deep-copied on enqueue, freed after send)
- ble_mgr.c: create the queue in tuya_ble_init, drain and release it
  in tuya_ble_deinit
- tuya_iot_dp.c: hand the ble obj/raw reports to the async queue

The ble path is only reachable while the cloud link is down, so this
is verified by build + elf symbols and no regression in cloud-online
tests.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant