Skip to content

Commit dff422b

Browse files
committed
Add closed-loop camera autofocus example
1 parent 5fb9fb3 commit dff422b

10 files changed

Lines changed: 803 additions & 0 deletions

‎Examples/autofocus/.gitignore‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,19 @@
1+
# Simulink build and cache output. Regenerated on demand, and slprj/ also holds the
2+
# accelerator-mode target, which is machine specific and must not be shared.
3+
slprj/
4+
*.slxc
5+
codegen/
6+
7+
# Autosaves and backups
8+
*.asv
9+
*.autosave
10+
*.slx.original
11+
*.mdl.original
12+
13+
# Compiled MEX
14+
*.mex*
15+
16+
# OS and OneDrive clutter
17+
desktop.ini
18+
Thumbs.db
19+
.DS_Store
174 KB
Binary file not shown.

‎Examples/autofocus/README.md‎

Lines changed: 50 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,50 @@
1+
# Closed-Loop Camera Autofocus in a Simulated Scene
2+
3+
A Simulink model of a camera that has to focus itself. An Unreal Engine parking lot is rendered
4+
to RGB and depth, a physically based defocus blur turns the depth into real optical blur, and a
5+
contrast-detect autofocus loop drives the focus distance using nothing but the blurred image.
6+
7+
![Autofocus running on the parking lot scene](autofocusCompare.gif)
8+
9+
Left: the camera output, with the region the autofocus measures marked in red. Right: the focus
10+
distance the loop commands, against the actual depth of the subject in that region.
11+
12+
## How it works
13+
14+
**Contrast is the signal.** The loop
15+
measures how much fine detail sits in a small region at the centre of the frame — a sharp image
16+
has more gradient energy than a blurred one — and that provides a target to optimize.
17+
18+
**An integrator climbs the hill.** The control loop dithers the focus by a tiny amount, sees which way contrast moved, and integrates that gradient, so the focus command
19+
is an accumulated output that simply holds still once it is on the peak.
20+
21+
**Stateflow implements a sweep.** An integrator can only climb the hill it is already on,
22+
and when the camera pans to view something at a different distance the old peak is gone. The model
23+
notices the drop in contrast, sweeps the focus across the range to find the new peak, and hands
24+
control back to the integrator to hold it.
25+
26+
## Requirements
27+
28+
Built and tested in **MATLAB R2026b**. You need:
29+
30+
- Simulink
31+
- Stateflow
32+
- Automated Driving Toolbox
33+
- Computer Vision Toolbox
34+
- Image Processing Toolbox
35+
36+
The 3D scene runs in the Unreal Engine co-simulation that ships with Automated Driving Toolbox.
37+
38+
39+
40+
## Files
41+
42+
| File | Role |
43+
| --- | --- |
44+
| `AutofocusScene.slx` | the model |
45+
| `exportAfVideo.m` | runs the model (or reuses the cache) and writes `autofocusCompare.gif` |
46+
| `helpers/blurSceneImage.m` | the optics the model uses, called by the `DefocusBlur` block |
47+
| `helpers/simulateDefocusBlur.m` | the defocus blur renderer |
48+
| `helpers/defocusBlurParams.m` | renderer defaults |
49+
| `helpers/nearWeightedDepth.m` | the reference depth statistic plotted on the right |
50+
8.21 MB
Loading

‎Examples/autofocus/exportAfVideo.m‎

Lines changed: 284 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,284 @@
1+
%exportAfVideo Export a side-by-side autofocus video: the blurred camera image with the
2+
% focus region marked, next to the focus command plotted against ground-truth depth.
3+
%
4+
% The point of putting these two panels in one frame is that neither is convincing alone.
5+
% The picture shows whether the subject looks sharp but says nothing about what the loop
6+
% was aiming at; the plot shows the command chasing a target but says nothing about
7+
% whether the result looks right. Side by side, a gap between the lines can be checked
8+
% against the picture immediately.
9+
%
10+
% The camera panel is composited at its native 640x480 rather than being drawn into a
11+
% figure axes. Letting a figure scale it to fit would resample the image and soften every
12+
% sharp region in it, which in a demo about defocus is exactly the wrong artifact to add.
13+
% Only the plot panel goes through a figure, and that one is rendered oversize and
14+
% downsampled, which makes its text cleaner rather than worse.
15+
%
16+
% Everything the video needs is cached to afVideoData.mat on the first run, so changing how
17+
% the video looks never costs another simulation: the sim is ~320 s, the render ~20 s. That
18+
% is also why the cache stores the raw depth PIXELS inside the focus region rather than a
19+
% summary of them - swapping median for mean, or for any percentile, then stays a code edit.
20+
% The crop is small enough to be free: 94x124x621 in single is 29 MB against 1.2 GB for full
21+
% depth. Delete the MAT file to force a fresh run.
22+
%
23+
% Writes autofocusCompare.gif next to this file.
24+
%
25+
% A GIF rather than an MP4 because the result is embedded in README.md, and a GIF plays
26+
% there with no player, no plugin and no click. The cost is that GIF is 256 colours with no
27+
% interframe compression, so both the frame count and the pixel count have to come down to
28+
% keep the file small enough to be worth embedding - see gifStride and gifScale below.
29+
%
30+
% The GIF lands in the code folder because it is the deliverable. The cache does not: it
31+
% carries every rendered frame at ~640 MB, so it stays in tempdir where it can be discarded
32+
% freely. Deleting it only costs the next run a simulation.
33+
34+
% Copyright 2026 The MathWorks, Inc.
35+
36+
m = "AutofocusScene";
37+
codeDir = fileparts(mfilename("fullpath"));
38+
cacheFile = fullfile(tempdir, "afVideoData.mat");
39+
outFile = fullfile(codeDir, "autofocusCompare.gif");
40+
41+
if isfile(cacheFile)
42+
fprintf("loading cached run from %s\n", cacheFile);
43+
S = load(cacheFile);
44+
else
45+
load_system(m);
46+
set_param(m, "SimulationMode", "accelerator");
47+
mws = get_param(m, "ModelWorkspace");
48+
spotFraction = mws.getVariable("AF_SpotFraction");
49+
50+
% name, block, output port
51+
probes = {"focus", "Autofocus", 1; "depth", "CameraRig", 2; "frames", "DefocusBlur", 1};
52+
ports = zeros(1, size(probes, 1));
53+
for k = 1:size(probes, 1)
54+
ph = get_param(m + "/" + probes{k, 2}, "PortHandles");
55+
ports(k) = ph.Outport(probes{k, 3});
56+
% Every sample, not decimated - this is a video, not a plot.
57+
set_param(ports(k), "DataLogging", "on", "DataLoggingLimitDataPoints", "off", ...
58+
"DataLoggingDecimateData", "off", "DataLoggingNameMode", "Custom", ...
59+
"DataLoggingName", char(probes{k, 1}));
60+
end
61+
62+
wall = tic;
63+
simOut = sim(m);
64+
elapsed = toc(wall);
65+
66+
for k = 1:numel(ports)
67+
set_param(ports(k), "DataLogging", "off", "DataLoggingNameMode", "SignalName");
68+
end
69+
save_system(m);
70+
71+
S = struct();
72+
S.Ts = mws.getVariable("Scene_Ts");
73+
S.spotFraction = spotFraction;
74+
% Total run, not just the pan: the camera keeps turning and tips down for Cam_TailTime after
75+
% the 100 deg sweep finishes, and the plot has to show that stretch rather than clip it.
76+
S.sweepTime = mws.getVariable("Cam_SweepTime");
77+
S.runTime = S.sweepTime + mws.getVariable("Cam_TailTime");
78+
S.focusMin = mws.getVariable("AF_FocusMin");
79+
S.focusMax = mws.getVariable("AF_FocusMax");
80+
S.focus = simOut.logsout.getElement("focus").Values;
81+
S.frames = simOut.logsout.getElement("frames").Values;
82+
83+
% Keep only the focus region, using the same mask the Sharpness and CentreDepth blocks
84+
% apply, so the numbers plotted describe the pixels the autofocus actually reads. Full
85+
% depth is 1.2 GB across the run; the crop is 14 MB.
86+
depth = simOut.logsout.getElement("depth").Values;
87+
[imH, imW, ~] = size(S.frames.Data(:, :, :, 1));
88+
rowsIn = find(abs(((1:imH)' - (imH + 1)/2)/imH) <= spotFraction/2);
89+
colsIn = find(abs(((1:imW) - (imW + 1)/2)/imW) <= spotFraction/2);
90+
S.spotDepth = single(depth.Data(rowsIn, colsIn, :));
91+
S.spotTime = depth.Time;
92+
93+
fprintf("%g s of scene time in %.0f s wall clock (%.1fx slower than real time)\n", ...
94+
S.runTime, elapsed, elapsed/S.runTime);
95+
% Uncompressed v7.3, because the frames dominate and compressing 440 MB of already
96+
% noisy render output costs far more time than it saves on disk.
97+
save(cacheFile, "-struct", "S", "-v7.3", "-nocompression");
98+
fprintf("cached to %s\n", cacheFile);
99+
end
100+
101+
t = S.focus.Time;
102+
% Caches written before the tail existed have no runTime field, and they are still wanted for A/B
103+
% comparisons, so fall back to the logged length rather than erroring on them.
104+
if ~isfield(S, "runTime")
105+
S.runTime = t(end);
106+
end
107+
d = S.frames.Data;
108+
n = size(d, 4);
109+
[imH, imW, ~] = size(d(:, :, :, 1));
110+
fprintf("%d frames of %dx%d, %.0f MB\n", n, imW, imH, numel(d)/2^20);
111+
112+
% Summarise the focus region, one number per frame: the median depth, weighted towards the near
113+
% content by 1/depth. See nearWeightedDepth for why it leans near and why a weight beats the flat
114+
% 25th percentile it replaced (0.35 m median difference from the focus command against 0.47 m, 83%
115+
% agreement within 20% against 75%). Changing this reduction never costs a run - the cache holds the
116+
% raw depth PIXELS, not a summary.
117+
%
118+
% A depth-histogram reference was also tried and dropped: the nearest peak holding at least a fifth
119+
% of the region, NaN when none does. It needed a bin width, a smoothing window and two thresholds to
120+
% do slightly worse than one weight. What it established is worth keeping: over the opening pan this
121+
% region holds eight surfaces, none above 15% of it, so no single summary of it can agree with the
122+
% focus command there - expect the two lines to disagree for the first ~3 s and read the picture
123+
% instead. The loop is not wrong there; because depth of field is symmetric in DIOPTERS, its far
124+
% focus leaves 47% of the region acceptably sharp against the 49% best available at any distance.
125+
flat = reshape(S.spotDepth, [], numel(S.spotTime));
126+
spotStat = nan(numel(S.spotTime), 1);
127+
for k = 1:numel(spotStat)
128+
spotStat(k) = nearWeightedDepth(double(flat(:, k)));
129+
end
130+
spotDepth = interp1(S.spotTime, spotStat, t, "previous", "extrap");
131+
132+
% The focus region, matching the mask in the Sharpness and CentreDepth blocks exactly: the
133+
% spot is AF_SpotFraction of the width by the same fraction of the HEIGHT, so it is not
134+
% square in pixels. Drawing a square crosshair would mark a region the autofocus is not
135+
% actually measuring, which is worse than drawing none at all.
136+
cx = (imW + 1)/2;
137+
cy = (imH + 1)/2;
138+
halfW = S.spotFraction/2*imW;
139+
halfH = S.spotFraction/2*imH;
140+
fprintf("focus region: %.0f x %.0f px centred at (%.1f, %.1f)\n", 2*halfW, 2*halfH, cx, cy);
141+
142+
% Corner brackets plus a centre cross with a gap, so the marker frames the region without
143+
% covering the detail the metric reads. Built once - the region never moves. Red, matching
144+
% the focus command in the plot, so the two panels read as one picture.
145+
brk = round(0.3*min(halfW, halfH));
146+
gap = round(0.25*min(halfH, halfW));
147+
armX = round(0.55*halfW);
148+
armY = round(0.55*halfH);
149+
xL = round(cx - halfW); xR = round(cx + halfW);
150+
yT = round(cy - halfH); yB = round(cy + halfH);
151+
bracketLines = [ ...
152+
xL yT xL+brk yT; xL yT xL yT+brk; ...
153+
xR yT xR-brk yT; xR yT xR yT+brk; ...
154+
xL yB xL+brk yB; xL yB xL yB-brk; ...
155+
xR yB xR-brk yB; xR yB xR yB-brk];
156+
crossLines = [ ...
157+
round(cx)-armX round(cy) round(cx)-gap round(cy); ...
158+
round(cx)+gap round(cy) round(cx)+armX round(cy); ...
159+
round(cx) round(cy)-armY round(cx) round(cy)-gap; ...
160+
round(cx) round(cy)+gap round(cx) round(cy)+armY];
161+
focusRed = [0.85 0.25 0.15];
162+
marker = [0.95 0.15 0.10];
163+
164+
% Plot panel. The animation is a cumulative reveal on fixed axes, so it does not need 621
165+
% figure renders - it needs two. Measured, `print -RGBImage` costs 613 ms per call: 94% of the
166+
% whole export, against 8 ms for everything drawn on the camera panel. So render the plot twice,
167+
% once empty and once complete, and per frame take the columns left of the current time from the
168+
% complete one and the rest from the empty one. Fixed axes make that pixel-exact, and it takes
169+
% the export from ~5.5 min to ~20 s.
170+
fig = figure("Color", "w", "Position", [80 80 imW imH], "Visible", "off");
171+
ax = axes(fig, "Position", [0.115 0.115 0.855 0.83]);
172+
hold(ax, "on");
173+
grid(ax, "on");
174+
box(ax, "on");
175+
% The focus stops, so a saturated command is distinguishable from a settled one.
176+
yline(ax, S.focusMin, ":", "Color", [0.6 0.6 0.6]);
177+
yline(ax, S.focusMax, ":", "Color", [0.6 0.6 0.6]);
178+
hDepth = plot(ax, nan, nan, "-", "Color", [0.15 0.45 0.80], "LineWidth", 1.8);
179+
hFocus = plot(ax, nan, nan, "-", "Color", focusRed, "LineWidth", 2.2);
180+
xlim(ax, [0 S.runTime]);
181+
% Mark where the pan ends and the downward-tipping tail begins, so a change in the focus
182+
% command's behaviour after that line is attributable rather than mysterious.
183+
if S.runTime > S.sweepTime
184+
xline(ax, S.sweepTime, "--", "Color", [0.45 0.45 0.45]);
185+
end
186+
ylim(ax, [0 S.focusMax + 3]);
187+
xlabel(ax, "time (s)");
188+
ylabel(ax, "distance (m)");
189+
title(ax, "focus command vs depth in the focus region");
190+
legend(ax, [hFocus hDepth], ["focus command", "near-weighted depth in region"], ...
191+
"Location", "northwest", "FontSize", 9);
192+
193+
% Empty render first. The legend draws from each line's colour and style rather than its data,
194+
% so it comes out identical in both renders and needs no special handling.
195+
axPos = ax.Position;
196+
xl = xlim(ax);
197+
yl = ylim(ax);
198+
% print -RGBImage works on an invisible figure, where getframe is unreliable. It honours display
199+
% scaling, so the result is oversize by an unknown factor (960x720 here) - resize to the panel
200+
% size rather than assuming it came back at figure size. The downsample anti-aliases the text,
201+
% which makes it cleaner rather than worse.
202+
panelEmpty = imresize(print(fig, "-RGBImage", "-r0"), [imH imW]);
203+
set(hDepth, "XData", t, "YData", spotDepth);
204+
set(hFocus, "XData", t, "YData", S.focus.Data);
205+
panelFull = imresize(print(fig, "-RGBImage", "-r0"), [imH imW]);
206+
close(fig);
207+
208+
% Data coordinates to panel pixels, for the reveal boundary and the moving marker. Exact for a
209+
% 2-D axes with no data aspect ratio constraint, which is what this is.
210+
toCol = @(tv) min(max(round((axPos(1) + (tv - xl(1))/(xl(2) - xl(1))*axPos(3))*imW), 1), imW);
211+
toRow = @(yv) min(max(round((1 - (axPos(2) + (yv - yl(1))/(yl(2) - yl(1))*axPos(4)))*imH), 1), imH);
212+
213+
% GIF sizing. GIF has no interframe compression and a 256 colour palette, so the file grows
214+
% with frames x pixels and both have to come down: every second frame at half size is a quarter
215+
% of the data. Playback stays real time, because GIF stores frame delays in hundredths of a
216+
% second and gifStride*S.Ts = 0.10 s is exact in that unit.
217+
gifStride = 2;
218+
gifScale = 0.5;
219+
keep = 1:gifStride:n;
220+
nGif = numel(keep);
221+
delayTime = gifStride*S.Ts;
222+
gifW = round(gifScale*2*imW);
223+
gifH = round(gifScale*imH);
224+
225+
% Downsampling the camera panel is allowed here even though the note above forbids resampling
226+
% it, and the distinction is worth stating. What is forbidden is letting a figure rescale the
227+
% image by an uncontrolled factor to fit an axes. A deliberate halving is scale consistent:
228+
% apparent blur measures 2.421% of the image width across 640x480 down to 240x180, so radius
229+
% and frame shrink together and relative sharpness survives.
230+
%
231+
% Annotations are not scale invariant. They are drawn at native size and downsampled with
232+
% everything else, so text, strokes and the focus dot are grossed up by 1/gifScale or they
233+
% arrive too thin and too small to read. The bracket GEOMETRY stays in native pixels, because
234+
% it has to keep marking the region the autofocus actually measures.
235+
ann = 1/gifScale;
236+
237+
comp = zeros(gifH, gifW, 3, nGif, "uint8");
238+
for f = 1:nGif
239+
[~, i] = min(abs(t - S.frames.Time(keep(f))));
240+
241+
left = insertShape(d(:, :, :, keep(f)), "line", bracketLines, "ShapeColor", marker, ...
242+
"LineWidth", round(3*ann));
243+
left = insertShape(left, "line", crossLines, "ShapeColor", marker, "LineWidth", round(2*ann));
244+
left = insertText(left, [8 8], sprintf("t %5.2f s focus %5.2f m depth %5.2f m", ...
245+
t(i), S.focus.Data(i), spotDepth(i)), "FontSize", round(14*ann), "BoxColor", "black", ...
246+
"BoxOpacity", 0.55, "TextColor", "white");
247+
248+
% Reveal the plot up to now, then mark the current focus. Opacity must be given explicitly:
249+
% insertShape defaults filled shapes to 0.6, which would make the marker translucent.
250+
col = toCol(t(i));
251+
right = panelEmpty;
252+
right(:, 1:col, :) = panelFull(:, 1:col, :);
253+
right = insertShape(right, "filled-circle", [col toRow(S.focus.Data(i)) 5*ann], ...
254+
"ShapeColor", focusRed, "Opacity", 1);
255+
256+
comp(:, :, :, f) = imresize([left right], [gifH gifW]);
257+
end
258+
259+
% One palette for the whole animation, sampled across the run so it covers the cars the camera
260+
% pans onto late as well as the ones it opens on. A per-frame palette makes the colours crawl
261+
% between frames and costs more bytes overall, because every frame then carries its own colour
262+
% table. No dithering: dither is high-frequency speckle, and in a demo about defocus it would
263+
% read as detail that is not there.
264+
sample = comp(:, :, :, round(linspace(1, nGif, 12)));
265+
[~, map] = rgb2ind(reshape(permute(sample, [1 2 4 3]), gifH, [], 3), 256, "nodither");
266+
267+
for f = 1:nGif
268+
indexed = rgb2ind(comp(:, :, :, f), map, "nodither");
269+
if f == 1
270+
imwrite(indexed, map, outFile, "gif", "LoopCount", Inf, "DelayTime", delayTime);
271+
else
272+
imwrite(indexed, map, outFile, "gif", "WriteMode", "append", "DelayTime", delayTime);
273+
end
274+
end
275+
276+
info = dir(outFile);
277+
fprintf("\nwrote %s\n", outFile);
278+
fprintf(" %d of %d frames at %g fps = %.1f s, which is real time\n", ...
279+
nGif, n, 1/delayTime, nGif*delayTime);
280+
fprintf(" %d x %d, %.1f MB\n", gifW, gifH, info.bytes/2^20);
281+
282+
err = abs(S.focus.Data - spotDepth);
283+
fprintf(" vs near-weighted depth in region: median |error| %.2f m, within 20%%: %.0f%%\n", ...
284+
median(err), 100*mean(err < 0.2*spotDepth));
Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,19 @@
1+
function blurred = blurSceneImage(rgb, depth, focusDistance)
2+
%blurSceneImage Defocus one AutofocusScene frame at a given focus distance.
3+
% blurred = blurSceneImage(rgb, depth, focusDistance) applies simulateDefocusBlur
4+
% to one frame of the camera output, focused at focusDistance metres.
5+
6+
arguments
7+
rgb (:,:,3) {mustBeNumeric, mustBeReal}
8+
depth (:,:) {mustBeNumeric, mustBeReal}
9+
focusDistance (1,1) double {mustBePositive} = 4.5
10+
end
11+
12+
cameraHorizontalFov = 60; % degrees, set by the Simulation 3D Camera block
13+
14+
params = defocusBlurParams();
15+
params.FocusDistance = focusDistance;
16+
params.FNumber = 0.28;
17+
params.FocalLengthPixels = (size(rgb, 2)/2)/tand(cameraHorizontalFov/2);
18+
blurred = simulateDefocusBlur(rgb, double(depth), params);
19+
end

0 commit comments

Comments
 (0)