Description
BOORU CHARS OPEN DATASET is an attempt to consolidate and arrange available character-centric anime/CG/game art.
It uses a localized format suited both for batch processing and visual estimation and contains JPG sample images
with reasonable quality and also tag and technical metadata.
**This release of BOORU CHARS consist of :**
- 1593429 sample images, mogrified to 1280px (1024px for 1х1)
* grouped into 18 volumes (directories) by aspect ratio and year
* zipped by 1000 images according to statistics similarity
* with verbose file naming "%website% - %id% - %copyright% ~ %characters% (%artist%)" and also with some tags in EXIF
- several tab-separated texts with metadata
* post/image info (for samples, ogiginals and from imfgeboard) 1.593.429 rows
* collected tag info with some addons 35.222.997 rows
* listing for 32 torrents total with 3.839.005 pictures (almost 5 ТБ) which brought dataset to reality
- an example of dataset usage for body parts detection and character assembling
* detector (notAI-tech NudeNet) results for 2 volumes and output compositions (DIY)
* ~4000 more or less interesting visualisations
- some descriptive insights into above
* readme_RU/EN with code examples
* several zipped Excels with analytic results and SQL-s
* [more and improved info](https://github.com/aperveyev/booru_processor) mostly russian
[Earlier version of this dataset (2019, 512px)](https://nyaa.si/view/1206322)
I want to believe that this half-terabyte of data worth more than chia coin mining.