Combines SAM's edge precision with DINOv2's semantic features and CLIP's language generalization to solve the spatial awareness gap in open-vocabulary segmentation. Previous approaches using CLIP alone lacked fine-grained spatial awareness, while DINO-augmented methods still struggled with precise edge perception. OVS-DINO aligns structural features across these vision foundation models with language guidance, advancing dense prediction beyond predefined category sets.