logo_test

Author	SHA1	Message	Date
Rick McEwen	ea6fcec9ce	Remove hybrid text+CLIP matching approach The hybrid approach combined OCR text recognition with CLIP embeddings to improve logo matching accuracy. After extensive testing, the approach was abandoned because: 1. OCR quality on small logo crops is unreliable 2. Text filtering rejected correct matches as often as wrong ones 3. Best hybrid result (57.1% precision) was similar to baseline (55.1%) 4. Recall dropped significantly (52.6% vs 59.6%) 5. Added complexity (EasyOCR dependency, extra parameters) wasn't justified Removed: - Hybrid matching methods from DetectLogosDETR class - Text extraction and similarity methods - Hybrid test scripts and text_recognition.py module - Hybrid-related CLI arguments from test_logo_detection.py The baseline multi-ref matching with 0.70 threshold remains the recommended approach for logo detection.	2026-01-08 12:48:39 -05:00
Rick McEwen	49f982611a	Add hybrid text+CLIP matching and image preprocessing Hybrid matching combines text recognition with CLIP similarity: - If reference logo has text and detection matches: lower CLIP threshold - If reference has text but detection doesn't match: higher threshold - If reference has no text: standard threshold Image preprocessing adds letterbox/stretch modes for CLIP input to preserve aspect ratio instead of center cropping. New files: - run_hybrid_test.sh: Test hybrid matching configurations - run_preprocess_test.sh: Compare preprocessing modes Changes to logo_detection_detr.py: - Add preprocess_mode parameter (default/letterbox/stretch) - Add set_text_detector() for hybrid matching - Add extract_text() using EasyOCR - Add compute_text_similarity() with fuzzy matching - Add find_best_match_hybrid() with tiered thresholds Changes to test_logo_detection.py: - Add --matching-method hybrid option - Add --preprocess-mode option - Add hybrid threshold arguments	2026-01-07 15:09:09 -05:00
Rick McEwen	6685af72d9	Add similarity distribution analysis for debugging embedding quality - Add --similarity-details flag to test_logo_detection.py - Track true positive, false positive, and missed detection similarities - Compute distribution statistics (min, max, mean, stddev, percentiles) - Analyze overlap between TP and FP distributions - Suggest optimal threshold based on data - Show per-detection breakdown with top-5 matches - Create analyze_similarity_distribution.sh wrapper script - Supports baseline, finetuned, or both models - Saves output to similarity_analysis/ directory	2026-01-05 13:39:20 -05:00
Rick McEwen	94db5bd40b	Add embedding model selection and comparison test scripts - Update DetectLogosDETR to support both CLIP and DINOv2 models - Rename clip_model parameter to embedding_model - Add model type detection for different embedding extraction - DINOv2 uses CLS token, CLIP uses get_image_features() - Add -e/--embedding-model argument to test_logo_detection.py - Include model name in file output header - Add run_threshold_tests.sh for testing various threshold/margin values - Add run_model_comparison.sh for comparing CLIP vs DINOv2 models	2026-01-02 12:05:27 -05:00
Rick McEwen	41c75356d9	Add --output-file option for clean results output - Add --output-file argument to test_logo_detection.py that appends only the results summary (no progress indicators) to specified file - Add write_results_to_file() with detailed header showing test type and method parameters - Update run_comparison_tests.sh to use --output-file instead of tee/redirection, keeping console output separate from file output	2025-12-31 17:42:52 -05:00
Rick McEwen	41bc0c701f	Add simple matching method as baseline for comparison tests - Add find_all_matches() method to DetectLogosDETR that returns all logos above similarity threshold without any rejection logic - Add --matching-method simple option to test script - Update run_comparison_tests.sh to include simple matching as Test 1 - Update documentation to describe simple matching method	2025-12-31 17:36:18 -05:00
Rick McEwen	197e007591	Add margin check to multi-ref matching to reduce false positives The multi-ref matching method was missing a margin check against other logos, causing excessive false positives. This fix adds: - margin parameter to find_best_match_multi_ref() that requires the best logo's score to exceed the second-best by a minimum margin - Test script now passes --margin to both matching methods - Updated documentation to reflect margin applies to both methods Also adds run_comparison_tests.sh to run all three matching methods and compare results.	2025-12-31 11:23:47 -05:00
Rick McEwen	ddccf653d2	Initial commit: Logo detection test framework Add DETR+CLIP based logo detection library and test framework: - DetectLogosDETR class for logo detection and matching - Test script with margin-based and multi-ref matching methods - Data preparation script for test database - Documentation for API usage and test methodology	2025-12-31 10:42:36 -05:00

8 Commits