Skip to content

Archive

Computer Vision

3 articles
Artificial Intelligence 19 Sep 2026 7 min read

Monocular Depth Estimation with MiDaS and DPT

A single RGB camera records image coordinates and color, not the physical distance from the lens to every visible surface. Monocular depth models infer the missing depth structure from visual evidence learned during training. That distinction matters when using MiDaS or DPT: a convincing depth map does not automatically mean that pixel values are distances in meters. MiDaS is an open-source project for robust monocular relative depth estimation. DPT, or Dense Prediction Transformer, is an architecture for dense prediction that has also been used as a backbone in MiDaS models. Both make single-camera depth estimation practical, but neither changes the geometric ambiguity inherent in one unconstrained RGB image.

Artificial Intelligence 12 Sep 2026 9 min read

Stabilize Image Model Predictions with Test-Time Augmentation

Stabilize Image Model Predictions with Test-Time Augmentation An image classifier can give slightly different answers when the same subject is cropped, mirrored, or resized in a way that preserves its meaning. If those transformations are valid for the task, relying on one view leaves useful evidence unused. Test-time augmentation (TTA) runs inference on several valid views of one input and combines their predictions into a final result. TTA is simple to describe, but safe use depends on details that are easy to miss. A transformation must preserve the target, structured outputs may need to be mapped back before aggregation, probability averaging can affect calibration, and every extra view consumes inference capacity.

Artificial Intelligence 08 Sep 2026 9 min read

Masked Autoencoders for Visual Representation Learning

Labeled image datasets are expensive to build, but unlabeled images are often plentiful. A useful pretraining strategy is therefore to create a learning signal from each image itself instead of asking a human to annotate it. A masked autoencoder (MAE) does this by hiding part of an image and training a model to reconstruct the missing content. The reconstruction task is not usually the final product. Its purpose is to make the encoder learn visual representations that can later support tasks such as classification or detection.