Urban scenes are often crowded and complex with many static and dynamic objects with diverse surface materials and patterns, making it a challenge for any vision model to understand and segment them perfectly.
Below is a comparative example of applying the Meta AI Segment Anything Model (SAM) on one of our footpath view imagery from Sydney with no further tuning versus our own semantic segmentation model part of our DeepWalk™ AI stack. The SAM works quite well in general but tends to over-segment without knowing the context obviously. This still has the potential to speed up often tedious manual image labeling/annotation tasks for domain-specific applications. While our DeepWalk™ semantic segmentation model is specifically designed, trained, and tuned to produce more meaningful and useful segments for mapping purposes.
Vision foundation models are as exciting as Large Language Models (LLMs), even more. Another example is the Microsoft Florence Foundation Model if you'd like to learn more.