LIVE HARDWARE OUTPUT STREAM
60.0 FPS · AXI4-Stream
⚙️ INSIDE THE SILICON HARDWARE
RTL Pipeline Active
CYCLE-BY-CYCLE SILICON DATAFLOW (LIVE WIRES & BUSES)
CLK: ⎍ TICK
Watch 8-bit pixels stream from camera into dual BRAM line buffers, latch into 3×3 shift registers, feed the parallel shift-add tree, and trigger the bounding box unit.
BUS EVENT: Cycle 1,048,576 · Streaming Raster Pixel (240, 160)
LATENCY: 1 Cycle / Pixel
STAGE 1: DUAL LINE BUFFER MEMORY (BRAM FIFO)
ON-CHIP SRAM
Stores previous 2 image rows so the chip never re-reads from external memory. Real pixel byte values shift left each cycle.
Row Y-2 (FIFO 1):
Row Y-1 (FIFO 0):
Row Y (Stream):
STAGE 2 & 3: 3×3 SLIDING WINDOW & PARALLEL MAC ENGINE
1 CYCLE LATENCY
p00120
p01134
p02142
p10118
p11210
p12145
p20115
p21128
p22150
Gx (Horizontal Gradient):
+84
Gy (Vertical Gradient):
-32
Magnitude |Gx| + |Gy|:
116
Threshold Comparator (> 65):
PASS (EDGE PIXEL)
STAGE 4: HARDWARE BOUNDING BOX COMPARATOR REGISTERS
ZERO LATENCY
X_MIN
142
X_MAX
206
Y_MIN
88
Y_MAX
152
Throughput Rate
1 Pixel / Cycle
No stalls or memory contention
Clock Frequency
150 MHz
1080p @ 60 FPS = 124.8 MHz pixel clk
FPGA BRAM Utilization
2 × 18Kb BRAMs
Zero external DDR DRAM access
Multipliers / DSP Units
0 DSP Blocks
Power-efficient binary shifts (p << 1)
📄 SYNTHESIZABLE VERILOG RTL SOURCE CODE
sobel_core.v (Top Processing Pipeline)
line_buffer_fifo.v (Dual Row BRAMs)
window_3x3.v (Sliding Shift Registers)
bbox_tracker.v (Bounding Box Min/Max)
Frequently Asked Questions
What is the latency of this hardware image processing pipeline?
The latency is exactly 2 image lines + 4 clock cycles. The hardware must buffer the first 2 horizontal scanlines before the 3×3 matrix is valid. Once filled, it streams 1 processed pixel on every single rising clock edge without stalling.
Why does the Sobel operator require 0 DSP multipliers?
The Sobel kernel contains coefficients [-1, 0, 1] and [-2, 0, 2]. Multiplying an 8-bit integer by 2 is accomplished in digital logic by shifting the bit vector left by 1 position (
p << 1), which requires zero transistors. The entire convolution is implemented with a lightweight 3-stage adder tree.