Real-Time Stereo Vision on the DE1-SoC
Fast stereo vision for hardware depth perception.
Introduction
Stereo vision is a technique for extracting a 3D scene from two 2D images. By setting up two cameras slightly offset from each other and pointed at the same picture, we can measure the difference in where features appear in the two images, a direct function of the depth of that feature. We demonstrate stereo vision running in real time on an FPGA.
Stereo vision is a powerful technique, but it is computationally intensive. Most algorithms use block matching: for each pixel in the reference image, select a block around it, and find the block in the other image which best matches it by sliding it across the frame. This process is expensive, each pixel of the image requires many calculations on many pixels at many positions. Basic block matching is embarrassingly parallel, but realizing this parallelism requires getting pixels from memory in parallel, and reusing data read from memory optimally.
These challenges align well with the DE1-SoC. The FPGA on the board has enough configurable M10K memory for each row of pixels to be stored in its own memory bank, allowing full columns of both frames to be accessed simultaneously. Most of the LUTs are utilized to implement the block matching algorithm, a sum of absolute differences (SAD) applied to the block. Finally, the DSPs are used for rectification, the process of transforming the warped camera image to a rectilinear projection.
Stereo vision is an incredibly practical method of getting depth information. Cameras are cheap and typically already present in places where stereo vision may be utilized, like autonomous vehicles and robots. While physical systems like LiDAR are also capable of extracting depth information, such systems are complex, expensive, and heavy. Stereo vision can be implemented wherever there are two conventional cameras.
High-Level Design
Our pipeline runs an endless loop of capture → rectify → match → display. Two analog cameras feed the DE1-SoC's Video-In subsystem, which writes interlaced PAL frames into an on-chip RAM that we replaced with a custom row-banked variant. A radial undistortion pass rewrites the captured frame in place, and a 12-way parallel SAD engine sweeps disparity hypotheses across the rectified pair, streaming the resulting depth map directly to the VGA framebuffer.
The Algorithm
Stereo Vision and Block Matching
Stereo vision works by determining disparities between the same features in two frames. The disparity of a feature is defined as the difference between its position in the reference image and its position in the other image. Disparity is inversely proportional to depth, scaled by the distance between the two cameras.
We create this disparity map by block-matching every pixel in the reference (left) image to a position in the other (right) image. Block matching selects a rectangular section of pixels centered on the pixel whose disparity we are trying to find in the reference image. Then, starting at the same position in the other image, we compare this block to every block in the other image in one direction up to a runtime-programmable maximum disparity. On each comparison, we calculate a difference between the reference and object block. When we have compared the reference block to every block up to a maximum disparity value, we pick the disparity that minimizes the difference between the reference and comparison block.
Sum of Absolute Differences (SAD)
Block matching relies on a way to calculate the difference between two different blocks of pixels. We experimented with many approaches, a Census transform with Hamming distance among them, but we found we got the best results in software testing using the simplest, sum of absolute differences (SAD):
This is, in essence, an L1 norm of the difference between the two blocks.
1-Directional Semi-Global Matching
Semi-Global Matching (SGM) is an algorithm used in computer vision to compute a dense
disparity map using a pair of rectified stereo images. The algorithm computes the
similarity of each pixel on one stereo image with each pixel in a subset of pixels from
the other stereo image. For each pixel with coordinates (x, y), the set of
candidates is selected as {(x̂, y) | x̂ ≥ x, x̂ ≤ x + D}, where D is the
maximum allowed disparity shift.
However, relying only on a pixel-matching algorithm produces a lot of incorrect matches, especially in textureless or monocolored areas. To combat this, SGM uses a cost function that applies penalties for having vastly different disparities for adjacent/close pixels. Generally the cost function to be minimized is of the form:
where D(p, dp) is the dissimilarity cost, how different the
pixels look, and R(p, dp, q, dq) is the smoothness
cost, how different the disparity-map pixels are in a given region. SGM attempts to
minimize this equation by finding the disparity that is most similar to the initial
coordinate while making sure that the final disparity-map pixel remains smooth with
respect to its surrounding pixels.
In this project we adapted the algorithm to better suit our needs.
Standard SGM reduces the 2D optimization above into a sum of 1D optimizations along
multiple scan paths, recursively accumulating cost along each direction. Full
multidirectional aggregation is great at smoothing but a poor fit for our project:
each path direction would need its own running cost array of size D in
memory, and any reverse-direction path would require a second copy of the cost volume
or a second pass over the frame. Our SRAM was already saturated
holding the rectified frame and ring buffer, so we collapsed the smoothness down to a
single 1D path along the horizontal left-to-right direction, the natural sweeping
order of the block matcher. The penalty cost minimized per pixel is then:
where R applies P1 when |dp − dp−1| = 1
and the larger P2 when it differs by more. This gives us most of the
smoothing benefit at a small fraction of the hardware cost.
Software Pathfinding
Before any of this landed in hardware, every algorithmic choice was first prototyped in Python and validated against captured stereo pairs. The Python work serves three roles at once: a parameter-sweep playground for finding the right operating point, a golden reference for cross-checking the C and RTL implementations, and the offline calibration tool that produces the lens-correction coefficients the FPGA actually consumes.
The three implementations share the same parameter names, window size, max disparity, vertical shift, block size, number of SAD units. When a Python experiment locks in a value, the C reference picks it up unchanged and the RTL constants follow. That alignment is what makes hardware debugging tractable: if the FPGA disparity map disagrees with the Python reference at the same parameter setting, the discrepancy has to be a real RTL issue, not a config mismatch.
Lens calibration
The calibration work walked from "no calibration, just an analytic fisheye model" to "real calibration baked into the FPGA as a small polynomial." The first attempt parameterised the two cameras' wide-angle lenses under several analytic projections, equidistant, equisolid, stereographic, rectilinear, and remapped the cropped active region with OpenCV, blending between identity and the model by a tunable correction strength. An autotune loop scored each candidate remap with edge-tracking and Hough-line "is this line straight after warp?" objectives. This never produced a usable stereo pair: the center looked corrected, the periphery didn't.
Switching to a real CharUco-board calibration with a fisheye
intrinsic model gave per-lens calibration matrices, which OpenCV's standard
rectification pipeline turned into a per-pixel gather map. That map was packed into
a 21-bit {valid, src_y, src_x} ROM format the RTL could consume
directly. The full LUT was too big to keep on-chip alongside the SAD engine and the
framebuffers, so a final pass compressed it into a small fixed-point polynomial, a
2D degree-6 fit in Q18, and a purely radial fit in Q15 with per-side optical-center
offsets. The radial form is what survives in the final RTL, because
it gives most of the quality at a tiny fraction of the DSP/LUT cost.
Each lens was calibrated independently, and the analog front end clips pixels asymmetrically off each side of every captured frame before the left/right split. The two cameras therefore have different optical-center offsets, and those offsets aren't symmetric about the midline. The final RTL reflects this with two completely separate 256-entry Q15 scale tables (one per camera) and four per-side optical-center constants, a destination-side center and a source-side center for each camera.
The two captures below are the live VGA output of the same scene, a person holding a tape measure horizontally across the camera frame, first with the raw, uncalibrated wide-angle lens, then with the per-lens calibration table driving the radial mapper.
Picking the matcher and operating point
The first family of disparity experiments was about finding the right matcher and operating point. OpenCV's BM and SGBM matchers (with P1 = 8·blockSize², P2 = 32·blockSize², 3-way SGBM mode) served as a golden reference, and we wrote a custom windowed-SAD from scratch with box-filter aggregation and a winner-take-all loop over disparity. Both L1 and L2 cost modes were swept.
The sweep driver walked max-disparities {24, 32, 48, 64, 96} against window sizes {5, 7, 9, 11, 13}, and a parallel sweep pushed the grid much further (window 3..33, max-disp up to 224) to see how flat the metric landscape was. The sweep result was a 9×9 window at max-disp 32, the smallest setting where the disparity map kept its structure on our specific captures, and that's the operating point baked into the C reference. The final synthesized RTL ultimately ships at 5×5 because the ALM and routing budget at 12 parallel SAD workers couldn't absorb a 9×9 worker (per-worker logic grows exponentially with the block radius); see Resource Utilization for that trade-off. The SGM penalty constants exposed to the HPS at runtime follow the same P1 = 8·blockSize² / P2 = 32·blockSize² formula OpenCV uses internally.
Cost metric and preprocessing
A second family of experiments changed the cost metric itself, because raw absolute-difference SAD struggled on low-texture and exposure-mismatched regions. A Census + Hamming path replaces each pixel with the sign-pattern of its neighbours and matches by box-filtered popcount-XOR distances, invariant to local brightness differences between the two cameras, which we definitely had. The sweep driver also ran a preprocessing grid over CLAHE on/off, median-blur kernel sizes, and a Gaussian high-pass before matching.
The conclusion was mixed: Census beat SAD on textureless surfaces but was worse on edges with real depth steps, and each preprocessing knob fixed one failure mode while introducing another. The RTL keeps SAD as the primary path and leaves Census as a parallel implementation we can swap in; CLAHE and median-blur stayed as runtime knobs on the C reference rather than burned into the RTL.
The 2D (d, vy) search
The third family was the 2D (d, vy) search, the cleanest example of a Python finding we couldn't get into the FPGA in time. OpenCV's SGBM assumes perfectly horizontal epipolar lines, and ours weren't: each lens was calibrated independently, the two cameras weren't perfectly co-planar, and remapping left sub-pixel row error that SGBM couldn't see through. So the Python and C matchers also try ±V row shifts of the right image for each candidate disparity and keep the (d, vy) pair with minimum aggregated cost. Real images need 1–3 rows of slack across the frame.
This was the largest-quality jump in the whole Python sweep, and it's a first-class
parameter in the C reference matcher. On the RTL side we never shipped it, the
row-banked on-chip RAM, the per-stripe SAD units, and the radial-mapper ring buffer
already filled the project's time budget, and a 2D (d, vy) search would
have needed the SAD engine to consume BLOCK_SIZE + 2·V rows of
right-image context at once instead of BLOCK_SIZE. So vshift stayed a
Python/C-only result that explains the residual row-misalignment artifacts we still
see in the FPGA disparity map.
Hardware Design
Our stereo vision system is built on the Terasic DE1-SoC. The Cyclone V FPGA fabric hosts a baseline system that provides a bridge to the ARM HPS, a Video-In subsystem that decodes the analog camera feeds into a single interlaced frame, and a VGA-out pixel DMA driving a framebuffer in SDRAM. To achieve real-time performance, we extended this baseline with three custom hardware subsystems that together form our depth pipeline:
- Row-Banked On-Chip Memory. A custom memory architecture that exposes a parallel read bus, allowing the SAD engine to fetch one byte from every row of the frame simultaneously.
- Radial Undistortion Pass. An in-place rectification pipeline that corrects lens distortion before matching begins.
- Parallel Block-Matching Engine. A 12-way parallel SAD compute unit that produces a dense disparity map, incorporating an SGM-inspired smoothness penalty.
These subsystems are sequenced by a four-state top-level finite state machine (FSM) that cycles through filling memory with rectified pixels, draining staging buffers, running the SAD compute, and reading back the results. Crucial parameters like window size, maximum disparity, and stripe height are configured at synthesis to match our software models, while SGM penalties are exposed to the HPS for runtime tuning.
Row-Banked On-Chip Memory
A key constraint shaped our entire architecture: even a small reference block compared against a wider search window requires fetching dozens of pixels per disparity hypothesis. A standard single-port memory caps throughput at one pixel per cycle, which would make the cost-per-disparity dozens of clock cycles.
To realize the embarrassingly parallel nature of SAD, we built a custom memory architecture that physically banks the frame along its Y-axis. Each of the 200 image rows is stored in its own dedicated memory block. While write access is multiplexed between the Video-In DMA and the HPS, we implemented a private, parallel read bus. This allows the hardware to fetch an entire column of pixels, one byte from every row, in a single clock cycle. This massive increase in read bandwidth is what makes the 12-way parallel SAD economically reachable.
Implementing this massive fan-in required specific architectural optimizations:
- Direct Memory Instantiation. Relying on standard inference for the memory blocks proved unreliable across different coding styles, leading to internal synthesis errors. We explicitly instantiated dual-port RAM primitives for each row to ensure stability.
- Two-Stage Read-Data Multiplexing. The 200-way multiplexer required for the read bus created a severe routing hotspot. We resolved this by splitting it into a registered 16-way intra-group multiplexer followed by a 13-way inter-group multiplexer. This added a cycle of latency but dropped placement and routing complexity by an order of magnitude.
- Global Write Arbiter. Rather than replicating write logic 200 times, a single arbiter broadcasts the winning write data and address to every row, with per-row logic gating only the byte enables. This saved over 3,000 ALMs of logic resources.
Rectification
Block matching only works if corresponding features lie on the same horizontal scanline in both images. Raw cameras with cheap wide-angle lenses don't give you that for free, straight lines in the world bend, and the two cameras' lenses don't match each other exactly. The analog front end on the DE1-SoC also clips about 40 pixels off each side of every captured frame before the left/right split, so each camera's true optical center doesn't sit at the geometric center of its cropped half. All of this has to be undone before SAD ever runs. Rectification is the geometric remapping that fixes it, and we bake the whole pass directly into the FPGA so the SAD engine never sees a raw pixel.
Because the two cameras' optical centers don't match and aren't symmetric about the split, the final RTL carries two completely independent per-camera calibrations: a separate radial scale table for each side, plus a separate destination-side optical center and source-side optical center per side. The pipeline selects which calibration to use based on which half of the frame the current destination pixel falls in. The offline calibration work that produces those tables and centers is described in the Software Pathfinding section.
The rest of this section walks the hardware pipeline in the order it runs: the radial mapper that converts a destination pixel into a source pixel, the staging ring that lets us rewrite the frame in place without corrupting our own input, and the four-phase scheduler that keeps the rectified frame consistent against an autonomous capture DMA.
The 11-stage Radial Mapper
The radial mapper is the hardware that consumes the per-camera calibration. It uses an 11-stage fixed-point pipeline: it consumes one destination pixel per cycle and emits the corresponding source coordinate eleven cycles later, calculating the vector to the image center, applying a radial coefficient, interpolating from the calibration lookup table, and applying center offsets.
The mapper is an eleven-state machine that takes one destination pixel
(dst_x, dst_y) and produces one source pixel (src_x, src_y)
over eleven clock cycles:
- ST_IDLE. Latch the destination coordinate, decide whether the pixel lives in the left or right half of the side-by-side frame, and load the four per-side calibration constants, destination lens center (Cx, Cy) and source lens center (Csrc, x, Csrc, y).
-
ST_VEC. Subtract the destination center to get a Q8 offset vector
(
dx, dy) pointing from the lens center to the pixel. - ST_R2. Square each component and sum them to get r², the squared radial distance from the lens center.
- ST_POS. Multiply r² by the constant 92 to produce a 32-bit Q16 LUT position, whose upper bits are the LUT index and whose lower sixteen bits are the interpolation fraction between adjacent table entries.
- ST_LUT. Read the two LUT entries flanking that index, s₀ = LUT[idx] and s₁ = LUT[idx+1], clamping to the last entry past the end of the table.
- ST_INTERP. Compute the difference (s₁ − s₀) and multiply it by the saved fraction, giving the partial slope-times-distance term of a linear interpolation.
- ST_SCALE. Add that interpolated correction back onto s₀ (with round-to-nearest) to finish the interpolation and produce a smooth Q15 scale factor for this exact radius.
-
ST_MUL. Multiply the original (
dx, dy) vector by that scale, stretching or shrinking it radially to the source-side geometry. - ST_OFFSET. Re-anchor the scaled vector on the source-side lens center by adding (Csrc, x, Csrc, y), producing the candidate source coordinate in Q8 fixed-point.
-
ST_CHECK. Round the Q8 coordinate to integer pixels, apply the
right-half offset of 319 if appropriate, bounds-check against the captured-frame
size, and drive the output registers
src_x,src_y, andvalid. -
ST_DONE. Pulse
done, dropbusy, and return toST_IDLE, ready to accept the next destination pixel.
Because this is a strictly sequential state machine rather than a pipelined one, the next pixel has to wait until the current eleven-cycle traversal completes before it can begin.
In-place rewriting and the staging ring
The mapper tells us which raw pixel to read for every rectified output pixel. The harder problem is where to put the result. Two hard constraints shape the rest of the pipeline:
- The on-chip SRAM is only about 200 KB, which is just barely enough to hold one 640×288 grayscale frame. There is not enough room for both a raw copy and a rectified copy, so the rectified frame has to overwrite the raw frame in the same memory, byte for byte.
- The Video-In subsystem's DMA is autonomous: it streams freshly captured PAL frames into the SRAM continuously on its own clock, without coordinating with anything else on the chip.
Every odd piece of the rectification pipeline is a direct consequence of one of these two facts.
The row-drift problem
Rectification itself is conceptually simple. For every output pixel
(dst_x, dst_y) we want to produce, the radial mapper tells us which raw
pixel (src_x, src_y) it came from in the captured frame, so we read that
raw byte out of SRAM and stash a rectified byte back into SRAM at the destination
position. The catch is that the lens distortion makes the mapper's source row drift by
up to about twelve rows above or below the destination row. So while
we are working on rectified row 60, the mapper might still ask us to read raw row 50
as a source pixel. If we had already overwritten raw row 50 with its rectified version
several rows earlier, we would now be reading our own rectified output as input, which
produces garbage.
The general rule is that raw row r is safe to overwrite only once every
later destination row that might still touch it is done reading from it. With a
maximum vertical drift of K = 12, that means raw row r is
safe only after destination row r + K is finished.
How the 13-slot ring solves it
This is where the undistort_ring comes in. Instead of writing each
rectified row straight back into SRAM, we stash it in a thirteen-slot circular buffer
for K more destination rows before committing it. When we finish
destination row 0 we put it in slot 0; row 1 into slot 1; and so on until row 12 fills
slot 12. Once we finish row 13 the ring is full, so before we put row 13 into slot 0
we first drain the old contents of slot 0 (rectified row 0) into SRAM row 0. At that
moment we are past row 13, the worst-case source row we could still need is row 0,
and we are already done with it, so raw row 0 is safe to overwrite. The ring is, in
essence, a K-row delay line that guarantees the safety margin holds for
every raw row.
The four-phase scheduler
The ring buffer is half the solution. The other half handles the autonomous DMA. The Video-In DMA, left to itself, would just keep writing fresh raw frames into the SRAM all the way through rectification and matching, randomly clobbering rectified rows we just produced and feeding the SAD matcher a mix of rectified and raw bytes. To stop that, the design holds a write-lock on the SRAM's s1 port for the entire rectify-then-match window, which makes the DMA's writes silently fail. That lock is essential for SAD to see a consistent frame, but it also means that while the lock is held, no new frames are getting captured.
So the top-level FSM serializes the whole system into four phases that take turns holding and releasing the lock:
- WAIT. The lock is released for ~50 ms, comfortably longer than one 40-ms PAL frame, so the autonomous Video-In DMA is guaranteed to lay down at least one fresh capture.
- FILL. The lock re-asserts. The radial mapper walks every destination pixel in raster order, raw bytes come out of SRAM, rectified bytes go into the ring, and the ring's oldest slot continuously drains back into SRAM in place.
- DRAIN. The last thirteen rectified rows still sitting in the ring after FILL ends are flushed to their matching SRAM rows, leaving SRAM holding a fully rectified frame.
- COMPUTE. The SRAM is still locked, and the parallel SAD engine sweeps the rectified frame, streaming disparities straight to the VGA framebuffer. When it completes, the cycle returns to WAIT.
Both the ring buffer and the WAIT phase are forced moves once you commit to fitting only one frame on chip and sharing that frame with an autonomous capture path. The full rectification subsystem, calibration coefficients, mapper FSM, ring buffer, and four-phase scheduler, exists so that the SAD engine in the next section can simply assume a clean, rectified frame is sitting in SRAM whenever COMPUTE fires.
SAD Block-Matching Engine
With a rectified frame sitting in the row-banked SRAM, the block matcher is the engine that actually computes disparities. The choice of SAD as the cost function, the max-disparity bound, and the SGM penalty formulation were all locked in during the Python sweep work, see the Software Pathfinding section for that pathfinding. The synthesized RTL block size (5×5) is set by the ALM budget rather than by the sweep result; see Resource Utilization for that trade-off.
The parallel-stripe architecture
The block matcher implements a parallel-stripe sweep over the rectified frames through several interconnected components:
- Striping. The 200-row frame is divided into 12 horizontal stripes. Each stripe is assigned its own dedicated SAD engine with a private sliding window. All 12 engines are clocked together, evaluating one disparity hypothesis at 12 different vertical positions simultaneously.
- Column Prefetch. A three-state handshake interface sits between the main engine and the SRAM. It translates stripe requests into parallel memory reads, latching full columns for both the reference and search images.
- Sliding Window. To optimally reuse data read from memory, each stripe utilizes a shift register file for its reference block. As the window shifts right across the image, it only needs to insert the freshly-prefetched column, keeping the rest of the block stable and eliminating redundant reads.
- Outer FSM. The control logic manages three nested loops: sweeping the disparity hypothesis up to the maximum limit, sliding the reference block along the X-axis, and shifting the windows down the Y-axis to cover the full stripe.
- 1-Directional SGM. We integrated our adapted SGM algorithm directly into the hardware. For each disparity hypothesis, the engine calculates a penalized cost: the raw SAD value plus a smoothness penalty. If the hypothesis differs from the adjacent pixel's best disparity by one, a small penalty is applied; if it differs by more, a larger penalty is applied. This single-pass approximation of SGM requires minimal hardware (a small adder and multiplexer per unit) but significantly smooths the output.
Cycle-by-cycle animation
The animation below walks through the mem_block_intf pipeline one clock
at a time, showing how the column prefetch, the per-stripe sliding windows, and the
SAD compute units stay in lockstep. Three units (red, orange, indigo) are highlighted
so you can watch how the same column read feeds multiple stripes in parallel.
HPS Software Design
The HPS software does not perform any load-bearing computational tasks. It is used as
the user interface, allowing the user to set parameters from the command line. The
command-line interface is implemented as a simple read-eval-print loop (REPL). The
user can set a parameter by typing the name of the parameter and value, e.g.
maxd 85 sets the maximum disparity value to 85. Updates to parameter
values are processed in real time. The HPS maps a PIO port for each of the
configurable parameters:
- maxd, the maximum disparity value which is checked.
- mind, the minimum disparity value which is checked.
- P1, the penalty assigned to a disparity value one away from the previous value.
- P2, the penalty assigned to a disparity value more than one away from the previous value.
- edge, enable the Altera Edge Detection IP.
Software Control Interface
Communication between the ARM processor and the FPGA is handled through Parallel I/O
(PIO) ports via the Lightweight HPS-to-FPGA bridge. The C program running on the ARM
processor maps the physical hardware addresses, including the bridge, VGA buffers,
and SRAM, into virtual memory using Linux's /dev/mem interface.
Because the FPGA-side PIOs are strictly output-only from the processor's perspective, the software maintains a local mirror of all tuning parameters. Upon booting, the program immediately initializes the hardware with a known-good default configuration, establishing baseline SGM penalties and disparity bounds so the stereo pipeline begins functioning immediately.
The REPL accepts terminal commands over SSH, allowing the user to adjust hardware parameters on the fly. Through this interface, users can dynamically tune the semi-global matching penalties, alter maximum and minimum disparity bounds, or toggle edge detection. The software sanitizes user input, such as clamping negative values, and pushes the updates to the hardware registers, altering the FPGA's behavior on the very next frame.
Frame Capture Utility
To evaluate the hardware's performance, we integrated a custom frame-capture utility
directly into the control software. When triggered via the REPL, the software reads
the SDRAM-backed VGA framebuffer and constructs a complete, uncompressed PNG
file entirely from scratch. We then copy these off the HPS using the
scp command.
Results
Speed
We were able to synthesize our design with 12 parallel SAD units executing simultaneously. At this level of parallel computation, we found that the framerate was ultimately bottlenecked by the undistortion pipeline. When the pipeline is full, it drops frames from the video input subsystem. In practice, with WAIT, un-pipelined rectification, and SAD compute running sequentially against the SRAM lock, the cycle takes about 150 ms per frame, ultimately limiting the framerate with rectification enabled to roughly 7 FPS.
With rectification disabled, we got worse results, since rectification is
necessary for horizontal sweeps through the frame to correspond exactly with
disparity changes. However, we did find a significant improvement in performance.
With N = 12, the frame was split into ceil(288/12) tall stripes, and
each worker had to do stripe_height * frame_width * max_disp + frame_width *
5 + stripe_height * 5 computations per frame. For our setup, this was
approximately 463,000 cycles. At the achieved clock frequency of 50 MHz, this
meant that a single frame was computed in 9.3 ms, corresponding with a frame rate
of about 108 FPS, significantly exceeding the 25 FPS framerate
of the PAL video standard we were using. In retrospect, dedicating more hardware
resources and design effort to speeding up the rectification pipeline would have
made computing faster than the 25 FPS PAL cameras achievable. However, this also
shows that our highly parallel block matching implementation is incredibly
performant, and would scale well to more frequent rectified video input.
Accuracy
We found that the stereo vision system performed the best at a distance of 1–3 m from the camera setup, with a camera distance of 6 cm. At this distance, it is possible to differentiate objects in the foreground from the background, with bodies and objects appearing as light, monochrome silhouettes against a dark, noisy background. The vision system is additionally able to pick up on gradients for objects which are angled towards or away from the camera, showing the change in the object's depth from the screen.
A known issue with basic block matching is that it performs poorly on monochromatic, untextured surfaces. This is because block matching assigns similar costs to every disparity within the surface, since they all appear identical without a priori knowledge of where that part of the block is within the surface. We ameliorate this using a simple, SGM-inspired penalty to SAD values to smooth out surfaces. This 1-way smoothing creates a noticeable difference between the left and right edges of objects, but it does significantly increase performance on monochrome surfaces, and larger objects in general, because it penalizes discontinuities in the disparity map.
To make the smoothing benefit concrete, we swept the SGM penalty configuration on the live FPGA pipeline against a fixed scene, a person holding a dark box in the lab. The top half of each VGA capture is the rectified stereo pair (left/right cameras), and the bottom half is the disparity map the matcher emits. Looking at the left half of each disparity map (the textured foreground) is the cleanest way to compare: the right half is dominated by the textureless wall beyond the box, where SAD has no signal to match against and even SGM can't recover the depth.
The takeaway matches what the cost function predicts: the SGM term is most
useful when both P1 and P2 are nonzero (panel 3), P1 alone or P2 alone leaves
the half it doesn't penalise free to drift, and pushing both magnitudes too high
(panel 4) starts erasing real depth structure. The values exposed at runtime
over the HPS REPL (P1, P2) let us pick the operating
point per scene without recompiling.
Holding the SGM penalties fixed (the panel-3 "both penalties" configuration above), we then walked toward the camera rig to see how silhouette quality scales with range. The five frames below are ordered from farthest on the left to nearest on the right. Disparity is inversely proportional to depth, so closer objects produce a larger pixel offset between the left and right cameras and therefore appear brighter in the disparity map; far objects collapse toward the noise floor.
The progression matches the depth sensitivity the geometry predicts: with a
6 cm baseline and max_disp = 32, the matcher's
useful range tops out around 1–3 m. Beyond that, even a real depth step shifts the
target by less than a pixel between the two cameras, the SAD cost surface
flattens out, and the disparity map collapses into the textured-background noise
regime visible in panel 1.
The same experiment in motion: continuous walks back and forth in front of the rig, captured as the disparity map updates frame-by-frame. The two clips below are the same scene under different operating points and concretely show the block-size ↔ worker-count trade-off the design analysis above discusses.
mind
raised over the default. The shipped configuration. The smaller block
gives a noisier silhouette per frame, but four times more workers run
the SAD compute much faster. Lifting the minimum disparity narrows the
useful depth range, so the same physical motion sweeps through more of
the output range, gradient changes look more dramatic as the subject
walks back.
Resource Utilization
We were able to successfully compile our design with a block size of 5 and 12 parallel SAD workers. We found that the limiting resource on the size of the design we could synthesize was the ALM usage needed to create larger or greater numbers of SAD compute units, together with the routing closure of the custom row-banked memory. Increasing the number of workers increases the ALM count roughly linearly, while increasing the block size increases the per-worker logic exponentially with respect to the radius of the block. A block size of 5×5 was the largest we could achieve while still having enough workers to not bottleneck the disparity map computation.
We initially struggled with M10K memory block usage because we inadvertently double-buffered incoming frame data. This was an artifact of two separate architectural decisions: the first was to put incoming pixel data into memory that was fragmented row-wise, the second was to store a rectified frame in that memory. As a result, we were initially using two separate memories for frame data: one was a monolithic on-chip SRAM synthesized from many M10Ks, which was written to by the video-in subsystem. Data from this memory was then rectified and copied to the segmented memory where it would be read in parallel by the SAD compute units. To resolve this redundancy, we wrote our own memory block, which exposed two interfaces: an Avalon MMIO interface to be written to by video input, and a parallel, asynchronous interface to be read from by the compute unit. This effectively halved our M10K utilization. Importantly, it allowed us to use a separate M10K unit for each row of each image, which allowed for maximum data parallelism. This memory block required many (13,268) ALMs, as the MMIO interface required muxing between 200 separate block RAMs which are physically spread across the die. This also led to very long compile times, as the fitter would struggle with this routing issue.
The actual Quartus fitter numbers below back up the architecture analysis above.
5CSEMA5F31C6: 24,079 / 32,070 ALMs (75%), 17,112 registers, 237 / 397 RAM blocks (60%), 17 / 87 DSPs (20%), and 1.88 / 4.07 Mbit of block memory used. The 75% ALM utilization is exactly what made adding more SAD workers a placement problem rather than a logic-availability problem.mem_block_intf consumes 13,268 ALMs on its own (the 200-way MMIO mux), and stereo_onchip_ram uses 200 M10Ks, exactly one per image row, the condition that makes the parallel column-fetch work. undistort_ring takes the remaining 13 M10Ks (the K = 12 + 1 staging-ring slots), and stereo_radial_mapper uses 6 of the 17 total DSPs for the Q15 multiplier chain. The four runtime PIOs at the bottom (pio_small_pen = P1, pio_big_pen = P2, pio_max_disp, pio_min_disp) are the HPS REPL parameters.Safety and Usability
The system runs entirely on low-voltage lab hardware, the DE1-SoC board plus two analog cameras and a VGA monitor, with no actuators, RF transmitters, lasers, or human-connected electronics. All parameter tuning is done over SSH through the REPL, and the disparity map is rendered live to the VGA output so anyone in the room can see whether a parameter change helped or hurt.
Conclusions
Our project demonstrated that fast block matching and stereo vision is achievable on the DE1-SoC using inexpensive PAL cameras. We were successfully able to demonstrate rectification, highly performant and parallel block matching, and a realistic depth map drawn to the VGA screen. Achieving this goal required solving a variety of problems. We had to measure and understand the geometry of the camera lenses we were using, utilizing OpenCV to calibrate our cameras. We tested a number of different block matching algorithms with different parameters and preprocessing techniques against our real-life camera setup, allowing us to thoroughly search the design space before implementing anything on the FPGA. It also required us to thoroughly understand the Cyclone V's capabilities, in particular the configuration of on-chip memory and utilization of the Avalon MMIO and streaming interfaces, enabling our blocks to talk directly to the Avalon video input IP. Finally, it required thorough consideration of the hardware architecture that was best suited for the task.
There are some architectural decisions that we made that either overcomplicated or constrained our design from the get-go that we would change for a future implementation. For starters, our implementation relies on storing an entire frame of video in memory, and then exploiting row parallelism to speed up computation. Because lines of video are received sequentially, a better approach would be to keep a much smaller line buffer of horizontal lines of video input, and parallelize across disparity sweeps rather than rows. This would have allowed us the same or better performance, as we would be able to exploit the same level of parallelism while using much less memory and ALMs.
Additionally, the greatest bottleneck on our video throughput was the rectification process. The latency for the un-pipelined rectification was 90 ms, significantly longer than the ~9 ms it took to compute the disparity map. Initially, we did not realize that this would even be necessary, and so we did not lay out a clear architecture for how the process should work. The clearest way to do the rectification would be to simply modify pixel coordinates in stream with a look-up table. We encountered two issues: one was that it was difficult to reverse the mapping from rectified coordinates to actual coordinates, and the second was memory usage for the look-up table. This combination of challenges led us to a massively overcomplicated multi-cycle rectification process that consumed many cycles of latency. This overcomplicated what should have been a relatively simple transformation that in principle could have been done serially in stream with a minimal impact on frame throughput. If we had anticipated that this would dominate the computation time, we would have spent much more design effort and been willing to spend more on-chip resources to simplify the pipeline.
Intellectual Property and Standards
We utilized some Altera IP in our design, in particular the video input IP, DMA controllers, VGA driver, and the memory-mapped IO standard. We used OpenCV for offline lens calibration and design-space exploration. We used iVerilog for RTL simulation and testing. We did not use any other code in our RTL or HPS code.
The system conforms to the standard PAL (25 FPS interlaced) video timing on the capture side and the DE1-SoC's 640×480 VGA timing on the display side. No custom standards work was required.
Appendix
A. Required Permissions
- We give permission for the report to be uploaded to the course website.
- We give permission for the demo video to be uploaded to the course YouTube channel.
B. Source Code Listings
Key files in the final build:
-
computer_15_640_video_mod/verilog/DE1_SoC_Computer.v, top-level integration. -
block_match.sv,block_match_pkg.sv, parameterized parallel SAD engine and shared parameter package. -
mem_block_intf_d.v, active row-banked memory interface (the_dvariant, not the older snapshot). -
column_prefetch_parallel.sv, parallel row-read column prefetch. -
stereo_onchip_ram.sv, custom Avalon-drop-in row-banked SRAM with side-door SAD conduit. -
stereo_radial_mapper_q15_pipe.v,radial_scale_lut_q15.vh, pipelined radial undistortion mapper and its calibration LUT. -
undistort_ring.sv,rectification_map.sv, staging ring and scheduler. -
stereo_c/, HPS REPL, frame-capture utility, NEON SAD reference. -
sterio-vision-python/stereo_depth_system/methods.py, Python matcher registry (SAD-L1/L2, Census+Hamming, vshift variants, OpenCV BM/SGBM golden). -
run_stereo_experiments.py, sweep driver. -
stereo_rtl_test/, Icarus Verilog image testbench.
C. Work Distribution and AI
We divided work for this project equally across members. Ezra contributed most significantly to the offline (software) design-space exploration, as well as the lens rectification and memory-banking RTL. Demetri contributed most significantly to the block matching RTL and top-level architecture. Daniel helped with RTL simulation.
We utilized generative AI tools to varying degrees for different parts of the project. For the block matching RTL, we wrote the majority of the RTL ourselves, and used generative AI tools (Claude Code, Microsoft Copilot) to identify bugs, or fix bugs that we had identified. This was useful for catching subtle bugs (e.g., we mistakenly wrote a shift register that did not shift the last index) and simplified the process of fixing more trivial bugs (e.g., small changes to state transition logic that we noticed in the waveform). For the segmented memory with Avalon MMIO integration, we used generative AI to tweak our RTL to conform to the Avalon interface standard, as well as to write a TCL script to get our module to appear in the Qsys GUI as a piece of IP that could be instantiated. The lens rectification RTL was more significantly AI-driven. On the software side, it generated most of the OpenCV code to extract lens parameters from a set of images, and on the hardware side we used it to generate a number of different implementations targeting different hardware utilization. We also used generative AI heavily for software design-space exploration: it allowed us to easily try out different block matching algorithms combined with different preprocessing steps and filtering. Finally, we used AI to generate some figures and interactive animations for this report, by giving it snippets of our RTL and asking it to create different kinds of displays (e.g., the interactive block matching animation).
We found that generative AI was in some ways a productivity multiplier, but in other ways made bad ideas tractable. AI was most effective for design-space exploration, because it allowed us to get a sense of what kinds of results were achievable, and how, very quickly. It also accelerated the process of writing RTL by catching and fixing bugs and writing useful testbenches. However, in some places it made overly complicated and poorly architected solutions possible, solutions we wouldn't attempt to write on our own. In particular, the RTL for the rectification pipeline, which was significantly AI-generated, was needlessly complex and served as a major performance bottleneck. Although the solution was correct, it was incredibly difficult to optimize and modify. On the other hand, our block matching RTL was almost entirely written by us, just utilizing AI to test and fix bugs. This RTL was comparatively very performant, easy to iterate on, and easy to understand conceptually. While the ability of generative AI to produce correct RTL is impressive, it tends to create solutions that are complex and difficult to iterate on without rigid guidance.
D. References
- Cornell ECE 5760 / U. Toronto Computer System 15 640×480 video-mod baseline.
- Terasic DE1-SoC user manual and Cyclone V device handbook.
- Intel/Altera Avalon and Qsys IP documentation.
- OpenCV stereo matching documentation: StereoBM, StereoSGBM, fisheye calibration.
- Hirschmüller, H. Stereo Processing by Semiglobal Matching and Mutual Information. IEEE TPAMI, 2008.
- Icarus Verilog and Pillow documentation.