How the FPGA SAD matcher operates on a 2D block window and runs
NUM_SAD_UNITS compute units in parallel — one per horizontal
stripe of the image. All units share the same FSM, the same memory request,
and the same disparity sweep, but each tracks its own best SAD.
Speed
slowCycle 0 / 0
Press Step or Play to begin.
Image view — 3 stripes processed in parallel
unit 0 stripe (rows 0–3) unit 1 stripe (rows 4–7) unit 2 stripe (rows 8–11) shared search window
Toy parameters:
BLOCK_SIZE = 3,
MAX_DISP = 5,
HALF_FRAME_WIDTH = 18,
FRAME_HEIGHT = 12,
NUM_SAD_UNITS = 3,
STRIPE_HEIGHT = 4.
The DE1-SoC build uses BLOCK_SIZE = 9, MAX_DISP = 85,
and FRAME_HEIGHT = 200 rows split across multiple parallel SAD units —
same algorithm, just more units and bigger blocks.
The synthetic stereo pair plants a different ground-truth disparity per stripe
(d=1,
d=2,
d=3) so each unit converges to a different
best_disp.
FSM & shared signals
state
IDLE
reg_phase (y in stripe)
0
reg_col_x
0
reg_disp
0
mem_req
0
mem_bank
—
mem_col (shared)
—
slide_reference
0
slide_matching
0
sad_compare_en
0
Per-unit SAD scoreboard
All units fetch the same mem_col every cycle (port B fans the byte out to all
rows in parallel). Each unit only reads its own BLOCK_SIZE rows and tracks its own best.
unit 0 · rows 0–3 · planted d = 1
SAD this candidate
—
best_sad
—
best_disp
—
unit 1 · rows 4–7 · planted d = 2
SAD this candidate
—
best_sad
—
best_disp
—
unit 2 · rows 8–11 · planted d = 3
SAD this candidate
—
best_sad
—
best_disp
—
Disparity stream (serialized)
disp_valid
0
emit_unit
—
disp_out_x
—
disp_out_y (= unit · STRIPE_HEIGHT + phase)
—
disp_out_value
—
How to read this
The image is divided into NUM_SAD_UNITS = 3 horizontal stripes of STRIPE_HEIGHT = 4 rows each. Every cycle, all three units operate on the same column at the same disparity step.
Each unit owns a BLOCK_SIZE × BLOCK_SIZE reference block at (col_x, g·STRIPE_HEIGHT + phase) on the left image, and a sliding match block at the same row offset on the right.
When mem_req fires, column_prefetch_parallel reads the requested column from every row of stereo_onchip_ram in one cycle (port B). Each unit indexes only its BLOCK_SIZE rows out of the returned bus.
The result is three independent best-SAD trackers running in lock-step. They tend to converge to different best_disp values because the underlying image content differs between stripes — in this synthetic demo we plant d = 1 / 2 / 3 respectively.
When disp exits its range, the FSM goes through INCR_X. It emits the three winners serially on the disparity stream — emit_unit cycles 0 → 1 → 2, one stream word per cycle (each acknowledged by disp_ack downstream).