Skip to content

About

Read and drive Windows GUI apps from a text-only agent: occlusion-free capture, OCR and vision for reading screens, pixel measurement, DPI-independent click coordinates, and a routine for screenshot evidence inside documents.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

windows-gui-vision

Notes and small tools for driving a Windows GUI from an agent that cannot see images.

I put this together while filling in a lab report that wanted a screenshot per item out of Proteus ISIS. The agent had no image input, screen captures kept grabbing whatever window was floating on top, and a good half of my clicks were landing on nothing at all. What is in here is what actually worked, together with the mistakes that cost me the most time.

The loop it is built around

capture the right pixels -> read them -> act -> prove the action landed

Capture. PrintWindow gets a window's own pixels, so occluding windows and the mouse cursor never appear and the target does not need focus. It works for GDI apps such as Proteus 7 and returns an all-black bitmap for GPU-composited clients. There is a foreground screen grab for that case, plus scripts that tell you which window you are looking at.

Read. A vision model for "what is this thing", Windows OCR for exact text with bounding boxes, and a handful of numpy scripts for anything that has to be measured rather than described: blob geometry, colour masks, ink profiles, before/after diffs.

Act. Keyboard first, accessibility indexes second, coordinates last. The coordinate case is handled by working in the application's own logical pixel grid and multiplying by a measured scale factor, so the same anchors survive a different monitor or DPI setting.

Prove it. Every canvas action commits silently. Diff two captures, or check for the ink you expected to appear. A click that returns success has told you nothing.

Quick start

python scripts/verify_env.py                        # interpreter, deps, key, OCR, smoke test
powershell -File scripts/wins.ps1 -Filter '*ISIS*'   # window handles, rects, titles
powershell -File scripts/capture_window.ps1 -OutPath shot.png -ProcessId <pid>
python scripts/calibrate.py --print shot.png --out calib.json
python scripts/see.py shot.png "which list row is selected?"

The only thing that needs a network connection and an API key is see.py. OCR and all the pixel tools run locally.

Install as a Codex skill

Copy the folder to %USERPROFILE%\.codex\skills\windows-gui-vision, or point skill-installer at this repository. SKILL.md is the entry point an agent reads; the references/ files are the detail it pulls in when it needs it.

Contents

Path What it is
SKILL.md the entry point: the loop, tool selection, rules of thumb
references/capture.md PrintWindow vs screen capture, occlusion, cursor, DPI, black captures
references/vision.md using a vision model, and where it lies to you
references/pixels.md the measurement scripts and when to reach for each
references/calibration.md logical pixels, scale factors, calibrate.py, layout.py
references/interaction.md clicking, typing, IME trouble, modal dialogs, canvas placement
references/coords.md the pointer readout, and turning a design coordinate into a screen pixel
references/docx-report.md turning captures into figures inside a Word document
references/proteus.md Proteus ISIS/ARES specifics, measured on a real install
references/dsn-format.md, references/dsn-generate.md editing Proteus .DSN files directly
references/dsn-append.md adding a component to a design from a script, verified, and what is still missing
references/dsn-build-circuit.md the workflow those pieces add up to: a circuit description into a design Isis opens, with the end-to-end run written up
references/dsn-templates.md lifting part records out of existing designs into a json library
references/dsn-wires.md adding a wire by script: the tail block, the link fields, and the verified recipe
scripts/ 30 helpers: capture, OCR, vision, pixels, calibration, layout, input, placement, probing, self-checks, .DSN reading, instances, wires, building a circuit from a description, and load checks

Everything under scripts/ is command line and prints plain text or json, so it composes in shell loops. SKILL.md and the reference files are what an agent reads; you can read them too.

Things it does not solve

  • GPU-composited windows. PrintWindow gives you black; the screen-grab fallback then depends on nothing being on top of the window.
  • The first calibration for a new app. Detectors find the canvas, an icon column and evenly spaced list rows, but naming which button is which is still a one-time manual pass.
  • Vision accuracy. It is fine at "what is this" and unreliable at coordinates, small text and anything that sounds like an inventory. The docs say which questions to avoid.
  • The geometry of a part you have just placed. Parts and wires can both be written by script now: adding an instance to a design that already embeds the device produces a file byte-identical to ISIS's own, and a wire can be written between coordinates, and a design built that way opens in ISIS. What is still measured rather than computed is where a freshly placed part's pins are, which is what a wire has to end on. references/proteus.md has both halves and the numbers.

License

MIT. See LICENSE.

About

Read and drive Windows GUI apps from a text-only agent: occlusion-free capture, OCR and vision for reading screens, pixel measurement, DPI-independent click coordinates, and a routine for screenshot evidence inside documents.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages