A semi-automatic evaluation system for LoRA fine-tuned diffusion models — generate, score, and analyze image quality across weight configurations in one pipeline.
Weights, seeds, base model
All (weight × prompt × seed)
Pairwise rating UI
Optimal weight + analysis
Accepts LoRA weights, evaluation metrics, base model, prompt lists, and seed lists. Iterates through all combinations and calls the inference engine automatically.
Web-based image grid with custom evaluation metrics. Supports image pair rating for efficient comparison and persistent score saving.
Score analysis with optimal weight calculation. Identifies representative best and worst prompts and images. Generates a final summary report.
Uses language models to help generate and refine evaluation criteria, reducing manual effort in defining quality metrics for specific LoRA styles.
Plots weight-quality curves across parameter ranges to find the optimal LoRA strength — balancing fidelity to the fine-tuned style against base model quality.
Evaluated with N=100 images in a user study. Users praised the concise weight-quality plots and representative pair comparisons.