Model or dataset
TheDesignFounder/DreamLayer-Eval avatar
TheDesignFounder/DreamLayer-Eval

DreamLayer Eval persists every run so you can replay it by run_id

DreamLayer Eval: open-source benchmarking for image and video diffusion models. Automate prompts, seeds, and metrics for reproducible results. A research project from DreamLayer AI (dreamlayer.io).

413 stars204 forksPythonGPL-3.0

At a glance

What is it?
A GPL-3.0 local benchmarking harness for image and video diffusion models, built around a React frontend, Flask services, SQLite run storage and a vendored ComfyUI. The design bet is that reproducibility comes from the storage schema first: prompt, seed, sampler, steps, CFG, model hash, LoRA stack and ControlNet config are all persisted so a run can be replayed.
Who is it for?
Adopt DreamLayer Eval if your team is already generating images and videos locally with ComfyUI and has been benchmarking by hand, because the stored run schema and the replay path are the parts you would otherwise build yourself.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 28 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Reproducibility is a storage schema, not a report you paste into a doc

Every run persists to SQLite carrying the prompt, the negative prompt, the seed, the sampler, the step count, CFG, a model hash, the LoRA stack, the ControlNet configuration, and all computed metrics. Once that is written, any run can be replayed by its run_id.

The model hash is the field that makes the rest useful. A seed plus a sampler plus a step count does not identify a model, so a run recorded without one cannot be reproduced even in principle, and the fact that it is stored as a hash rather than a filename means the record survives a file being moved.

The LoRA stack and ControlNet config are stored the same way, as resolved configuration rather than as a pointer to a preset. That is what makes the run_id replay meaningful after you have edited a workflow.

N prompts by M seeds by K samplers is the sweep that replaces the manual week

The headline claim is arithmetic rather than algorithmic. One run sweeps N prompts by M seeds by K samplers, and metrics compute live during generation rather than in a separate pass afterwards. A benchmark that would take one to two weeks manually is stated to finish in three to five hours per model.

The mechanism that makes that affordable is the metric cache. Metrics persist per run in a dedicated SQLite table and return instantly on re-fetch, and batch backfill endpoints recompute only what is missing across the full history rather than starting over.

That cache is also what makes runs comparable over time. If a metric is cached, adding a new prompt to the sweep does not recompute the old rows, so the earlier runs stay in the same units and the comparison remains honest.

A model appears in the dropdown only once its key is in .env

Weights do not ship with the repository, which keeps the download small. There are two ways to add models and they behave differently.

The API path is a .env file in the repository root with four keys: OPENAI_API_KEY, BFL_API_KEY, IDEOGRAM_API_KEY and STABILITY_API_KEY. The visibility rule is the part worth knowing: once a key is present the model appears in the dropdown, and with no key the feature stays hidden rather than appearing greyed out. So an empty dropdown is not a bug, it is the absence of a credential.

The offline path is the conventional one. Download .safetensors or .ckpt files from Hugging Face, Civitai or your own training runs, then drop them into Checkpoints/, Lora/, ControlNet/ or VAE/, which are created on first run. Settings then Refresh Model List, and the models show up. The tip for anyone keeping weights on a separate drive is to use symbolic links.

Two environment variables, and one of them defaults to a path in /tmp

The configuration surface is small and documented per script rather than globally. install_dependencies_linux reads DLVENV_PATH as the preferred path to the Python virtual environment, and its default is /tmp/dlvenv.

That default is the one to change on a machine that cleans /tmp on reboot, since a virtual environment living there is a virtual environment you rebuild for no reason.

start_dream_layer reads DREAMLAYER_COMFYUI_CPU_MODE, which is the escape hatch for machines without NVIDIA drivers and defaults to false. Everything else, from the API keys to the model directories, is handled by the scripts and the UI rather than by environment variables, which is a reasonable line to draw for a tool whose audience is researchers rather than cluster operators.

Two servers on two ports, and the ComfyUI one is vendored in the tree

Once started, the frontend answers on http://localhost:8080 and ComfyUI answers on http://localhost:8188. The application is a React frontend over Flask-based services with SQLite for run storage, and image work goes through ComfyUI.

Unusually for a Python project, ComfyUI is not a dependency you install, it is a directory at the repository root alongside dream_layer_backend/, dream_layer_frontend/, scripts/, tests/ and docs/. That is convenient for reproducibility and awkward for upgrades, since you are carrying a copy of another project and have to reconcile it yourself when ComfyUI moves.

The tree also keeps both start_dream_layer.sh and start_dream_layer_linux.sh, plus a .bat for Windows, so there is more than one way to launch the same thing and the README only names the generic one.

Reference-based metrics need a reference clip before they produce anything

The metrics split along one axis: whether they need a ground truth. CLIPScore, aesthetics, YOLO composition, temporal flickering and DINO consistency all work with no reference at all, which is why the harness can be useful on a checkpoint you have no baseline for.

SSIM, PSNR and LPIPS only run once you add a reference video, and FID operates on a reference set rather than on a single file. For FID that means the CIFAR-10 reference data has to be present first, which is what the dataset fetch step is for:

bash
python scripts/fetch_datasets.py

Object detection is the exception that needs no preparation: the YOLO model, yolov8n.pt at roughly 6MB, downloads itself on first use.

The image set is CLIPScore with ViT-L/14, FID, LAION aesthetic, color harmony, sharpness and YOLOv8 composition F1. The video set is FVD with I3D, SSIM, PSNR, LPIPS, temporal flickering, subject and background consistency via DINO, and motion smoothness. Custom metrics plug in, and audio benchmarking is described as on the roadmap rather than available.

The Cursor path promises five to ten minutes and admits macOS needs retries

The documented easiest route is not a command line. You download the repository, open the folder in Cursor, type run it or press the Run button, and Cursor walks through setup: installing Python and Node dependencies, creating a virtual environment, starting the backend and the frontend, and printing the localhost:8080 link. The stated time is five to ten minutes with no terminal needed.

The caveat is stated in the same breath: on macOS the PyTorch setup may take a few retries and you should keep pressing Run when prompted.

The manual scripts exist for all three platforms, one shell script each for Linux and macOS and a PowerShell one for Windows. The Windows path is the only one with a policy wrinkle, since it asks you to relax execution policy for the session first:

GPL-3.0, no tagged releases, and a codebase that carries several experiments

Licensing is GPL-3.0, which matters if you intend to embed this in something you distribute under other terms. The repository is not archived and the last push was 2026-09-09, but there are no GitHub releases, so there is no version to pin and nothing to diff a regression against.

The tree also reveals what this is mid-way through. GEMINI_INTEGRATION_DEMO.md, GENERATION_HISTORY_IMPLEMENTATION.md, test_gemini_integration.py, KaggleDreamLayer.ipynb, comfy-easy-grids/ and ui_mockup_grid_export.html are all sitting at the top level next to the application directories, along with an mkdocs.yml and a .claude/ directory.

Design notes sitting beside the code is normal for a research project and is a signal that the framework is still moving. The status line in the documentation says now live, which is a statement about the product page rather than about the release history.

Editorial conclusion

Adopt DreamLayer Eval if your team is already generating images and videos locally with ComfyUI and has been benchmarking by hand, because the stored run schema and the replay path are the parts you would otherwise build yourself. Skip it if you need audio metrics, which are stated as roadmap rather than shipped, or if you want a turnkey product, since the quickest documented route still installs Python and Node dependencies, builds a virtual environment and starts two servers. Before you commit, check the last push of 2026-09-09 against your own plans, confirm that the vendored ComfyUI directory at the repository root matches the version your workflow expects, and read the licence terms, because GPL-3.0 has consequences if you intend to wrap it in something you ship closed.

Frequently asked questions

What does DreamLayer Eval actually measure?

Image metrics are CLIPScore with ViT-L/14, FID, LAION aesthetic, color harmony, sharpness and YOLOv8 composition F1. Video metrics are FVD with I3D, SSIM, PSNR, LPIPS, temporal flickering, subject and background consistency via DINO, and motion smoothness. Custom metrics are pluggable and audio benchmarking is described as being on the roadmap.

How do I replay an earlier DreamLayer Eval run?

Every run is persisted to SQLite with the prompt, negative prompt, seed, sampler, steps, CFG, a model hash, the LoRA stack, the ControlNet configuration and all computed metrics, and any run can be replayed by its run_id. Metrics are also cached per run in a dedicated table so re-fetching returns instantly rather than recomputing.

Why does a model not appear in the DreamLayer Eval dropdown?

For API models, the corresponding key has to be present in the .env file in the repository root, one of OPENAI_API_KEY, BFL_API_KEY, IDEOGRAM_API_KEY or STABILITY_API_KEY. With no key the feature stays hidden rather than appearing disabled. For local checkpoints, the files go into Checkpoints/, Lora/, ControlNet/ or VAE/ and you then use Settings and Refresh Model List.

Does DreamLayer Eval run without an NVIDIA GPU?

Yes. The start script reads DREAMLAYER_COMFYUI_CPU_MODE, which is described as the setting to use when no NVIDIA drivers are available and defaults to false. The frontend serves on localhost:8080 and ComfyUI on localhost:8188.

What is the difference between DreamLayer Eval and DreamLayer AI?

They are stated to be two separate products. DreamLayer Eval is this repository, a free and open-source toolkit for evaluating and comparing image and video generation models. DreamLayer AI is a separate agent for image generation and editing, usable on the web or through its API, CLI or MCP server with tools such as Claude Code, Cursor and Codex.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. Project website
  4. README
  5. TheDesignFounder/DreamLayer-Eval on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/thedesignfounder-dreamlayer-eval.svg)](https://hysenlabs.com/projects/thedesignfounder-dreamlayer-eval)