This repository contains code for Fly Behavior startup project in collaboration with Fumika Hamada. The project is about studying the temperature preferences of fruit flies (Drosophila) over a 24-hour period, in order to better understand which genes mediate circadian rhythm. The goal of our collaboration is automate the process of counting fruit flies and estimating their temperatures in photos from experiments.
Links:
Contents:
The fly detection tools assume a standard format for data sets:
YYYY-MM-DD_apparatus/ the data set directory
├── photos/ original photos, in JPEG format
└── temperatures.xlsx temperature spreadsheet, in XLSX format
The data set directory must be named with the date and apparatus name. The
photos in photos/ can have any name but must have extension .jpg or
.jpeg. The .xlsx file can have any name. In general, it's not a good idea
to put spaces in file names (use at your own risk!).
For example, one of the data sets we used during development had this format:
2023-07-19_biden
├── photos/
│ ├── 2023-07-19_biden_photos_01.JPG
│ ├── 2023-07-19_biden_photos_02.JPG
│ ├── 2023-07-19_biden_photos_03.JPG
│ ├── 2023-07-19_biden_photos_04.JPG
│ ├── ...
│ └── 2023-07-19_biden_photos_42.JPG
└── 230719_Biden_Leia_template.xlsx
For an appropriately formatted data set, the fly detection workflow consists of two steps:
-
Arena Registration (the
registercommand). For each image in a data set'sphotos/subdirectory:- Standardize the brightness and contrast.
- Find the registration marks.
- Rotate as needed so that the image is not upside-down or sideways.
- Correct for perspective distortion so that the registration marks form a perfect rectangle.
- Crop to the rectangle.
- Save the cropped image in the data set's
arenas/subdirectory.
-
Fly Detection (the
detectcommand). For each image in a data set'sarenas/subdirectory:- Use the model to predict fly locations as bounding boxes.
- Remove bounding boxes with too much overlap (the model typically makes some redundant predictions).
- Estimate the temperature at each fly's location.
- Save the predicted fly locations and estimated temperatures to a
.csvor.parquetfile in the data set directory.
These steps are implemented as two different command-line commands (register
and detect, respectively).
All of the command-line commands have only one argument: a path to a TOML
configuration file. TOML is a plain-text configuration file format designed to
be easy to read and write. An example configuration file is provided in this
repo at configs/defaults.toml. A copy of the file with long-form
documentation is provided in this repo at
configs/defaults-long-comments.toml. You can open and edit
TOML files with a text editor such as Notepad++, TextEdit, or nano.
The TOML config file contains settings for each command, as well as physical
measurements (in centimeters) for each apparatus. We recommend that you create
a new TOML config file for each data set, so that you have a record of the
settings you used to process each data set. The easiest way to do this is to
copy configs/defaults.toml or another config file and then edit as needed.
In the TOML config file, the most important setting is data_path, which
should be set to the path to the data set directory. This setting and others
are documented in configs/defaults.toml.
Once you've created a TOML config file, for example
configs/2023-07-19_biden.toml, you can run arena registration. In a terminal,
navigate to the repo (with cd) and make sure the fly environment is
activated (mamba activate fly). Then run:
python -m src register configs/2023-07-19_biden.tomlThis will create an arenas/ subdirectory in the data set directory. You can
inspect the images in arenas/ to check that the arenas were detected
correctly.
Arena detection typically fails if the registration marks are covered in the
original photo or the original photo has poor brightness or contrast. As a
failsafe, you can manually specify the pixel coordinates of the apparatus
corners in the TOML configuration file with the arena setting. An example of
this is provided in configs/2023-08-07_biden.toml.
Next, you can run fly detection. In the terminal, run:
python -m src detect configs/2023-07-19_biden.tomlThis will create a predictions.csv file and predictions/ subdirectory in
the data set directory. The predictions/ subdirectory contains visualizations
of the predicted flies as JPEG images (one for each arena image). The images
show each detected fly's bounding box, an identification number (ID) for the
box, and grid lines every 0.5 degrees Celsius.
Note that box IDs is not linked across images, so box 1 for the first image
in a data set does not necessarily enclose the same fly as box 1 for the
second image. Box id is only provided as a way to easily remove incorrect
boxes.
The format of predictions.csv is described in the next section.
The predictions.csv file created by the detect command has one row for each
detected fly. The model detects a bounding box around each fly, so the columns
contain data about the bounding box. The columns are:
| Column | Description |
|---|---|
id |
identification number for the box (within the image) |
x_px |
x-coordinate of the box center, in pixels |
y_px |
y-coordinate of the box center, in pixels |
width_px |
width of the box, in pixels |
height_px |
height of the box, in pixels |
confidence |
confidence score for the box (from 0 lowest confidence to 1 highest confidence) |
path |
file path to arena image |
arena_width_px |
width of the arena, in pixels |
arena_height_px |
height of the arena, in pixels |
x_cm |
x-coordinate of the box center, in centimeters |
y_cm |
y-coordinate of the box center, in centimeters |
temperature |
estimated temperature at the box center, in degrees Celsius |
The file is a comma-separated values (CSV) file, which can be read and analyzed with data analysis software such as Excel, Tableau, Python, and R.
Note
The TOML config file also provides a setting to save the file in Parquet format. Parquet is an open-standard for data exchange that provides several benefits over CSV files.
The directories and files in this repository are:
configs/ TOML configurations for commands
data/ Data sets (files > 1MB go on Google Drive)
models/ Deep learning models for fly detection
notebooks/ Jupyter notebook source files (exploratory code)
src/ Python source code
.gitignore Settings file for git
README.md This file
fly.yml Main Conda environment (with OpenCV, etc)
fly-dev.yml Conda environment for development
fly-train.yml Conda environment for training the model.
tess.yml Conda environment for Tesseract
Each .md file in notebooks/ and .py file in src/ has a brief
description at the top of the file. The data/ and models/ directories are
not included with the repo, but typically need to be created to use the tools.
The fly detection tools can run on macOS, Windows, and Linux. They were developed and tested on macOS and Linux.
To get started, download a copy of this repository from GitHub to your computer. To do this, navigate to the repo's main page, click on the green "Code" button, and select "Download ZIP" from the menu. Pay attention to where you save the file. Once the download is complete, unzip the file.
Note
If you plan to edit the code, we recommend that you use git to download a copy of the repo instead. You can learn more about git from these DataLab workshop readers:
If you've configured git to connect to GitHub, you can use git clone to download a copy of this repository to your computer:
git clone git@github.com:datalab-dev/2023_project_hamada_fly.gitThe fly detection tools provide a command line interface. Familiarity with the Unix command line will make it easier to follow the instructions in this README and to use the tools. You can learn more about the command line from DataLab's "Introduction to the Unix Command Line workshop reader. To access the Unix command line:
- macOS: use the built-in "Terminal" application.
- Windows: install git and use the "Git Bash" application. The built-in "CMD.exe" application does NOT provide a Unix command line.
- Linux: use any terminal application (for example, xterm).
In your terminal, change directories to where you downloaded and unzipped the repo. Then change directories to the repo directory:
cd 2023_project_hamada_fly/Next, create a models/ subdirectory to store model files:
mkdir modelsGo to the Google Drive and download the file
models/2023-08-25_fly-detection.onnx to the models/ subdirectory you just
created.
We recommend that you also create a data/ subdirectory to store data sets:
mkdir dataTo install Python and the Python packages necessary for the for the fly
detection tools to run, we recommend using mamba. You can install mamba by
installing miniforge (formerly known as "mambaforge"). If you already have
conda installed, you can use that instead by replacing mamba with conda in
the commands below. You can learn more about these tools from this
section of DataLab's "Making Python Projects & Environments
Reproducible" workshop reader.
To recreate the Python environment required by the fly detection tools, run:
mamba env create --file fly.ymlThis will create an environment named fly. Finally, activate the environment:
mamba activate flyNow you're ready to use the fly detection tools!
Important
This section is about how to contribute notebooks to the repo.
Jupyter notebooks are stored in the repo in Markdown format (.md) via
Jupytext. This makes it easier to see changes to the notebooks in version
control and also avoids committing large images to the repo.
The conda environments in the repo include Jupytext. Make sure one is installed and active before running the commands below.
Whenever you create a new Jupyter notebook, say notebook.ipynb, run this
command to make it a paired notebook and generate a corresponding
notebook.md file:
jupytext --set-formats 'ipynb,md' notebook.ipynbThe notebook.md file is the one to commit to the repo. Once a notebook is
paired, the two files will automatically be kept in sync as long as you run
Jupyter in an environment that has Jupytext installed.
If you over need to manually sync a paired notebook, the command is:
jupytext --sync notebook.ipynbBe careful when doing this, because if you've somehow ended up with changes to
both notebook.ipynb and notebook.md, the older changes will be overwritten!
You generally won't need to run these if you've followed the instructions above.
To convert all .ipynb files to .md, run this command in the notebooks/
directory:
jupytext --to md *.ipynbYou can replace *.ipynb with a specific file name if you only want to convert
one file.
To convert .md to .ipynb, run this command:
jupytext --to ipynb *.mdSee the Jupytext CLI docs for more info. There is also a Jupytext extension for JupyterLab that can handle this process automatically.
Important
The model has already been trained, so it is not necessary to train the model in order to use the fly detection tools described above. That said, if new annotated data becomes available, additional training may improve accuracy.
The model is a You Only Look Once (YOLO) v8 object detection model. It was trained on the UC Davis Farm Cluster. The node used has 2 AMD EPYC 7713 64-core CPUs, 1TB RAM, and a NVIDIA A100 GPU. Training for 200 epochs took approximately 2 hours.
We created training data by template matching 15-degree rotations of a manually cropped fly to each image in these data sets:
2023-07-12_biden2023-07-19_biden2023-07-19_shiny2023-07-22_shiny2023-07-23_skywalker2023-07-26_shiny2023-08-09_biden
We then manually corrected these annotations (adding boxes around undetected flies and removing incorrect boxes) with the free MakeSense annotation software. This process is time-consuming, as template matching is often inaccurate. Going forward, it is likely more efficient to manually-correct predictions from the current fly detection model.
Before running template matching, create an outputs/ subdirectory in the
repo, then go to the Google Drive and download the file
outputs/fly_template.npz to the subdirectory you just created.
The command to run template matching is:
python -m src detect configs/defaults.tomlAs with other commands, the TOML config file contains settings such as which
data set to use. The resulting annotations are saved to the match_labels/
subdirectory of the data set directory.
After manually correcting the annotations in MakeSense, export the
annotations in YOLO format and unzip the resulting .zip file into a labels/
subdirectory of the data set directory.
You can prepare a training data set, combining annotations for several data sets, with the command:
python -m src assemble configs/train.tomlSee the TOML config file for settings.
Finally, to train the model, run:
python -m src.train configs/train.tomlWhen training is complete, the model will be saved in the models/ directory
of the repo. Note that the training script was designed and tested for training
the original pretrained YOLOv8x model; it was not tested for additional
training of the fly detection model, so some editing may be necessary.