ObjectDetectionPipeline
pipe(image, threshold=0.5) -> [{"score", "label", "box": {"xmin", "ymin", "xmax", "ymax"}}] in image pixels; an open-vocabulary model takes
candidate_labels: a list of labels, each one phrase, or the prompt
text itself in the family's comma-separated phrase grammar
("cat,remote control"; a period is part of a phrase, not a
separator).