Datasets
Data that companies have brought to Elixr. Seven datasets in total. Each sheet shows the classes of record inside with the text abstracted: you see the structure, never the content. Region height means something only where every class carries a measured figure, and the sheet says when it does.
- Mixpanel, Firebase
- Drive, Dropbox, Notion
Classes present
Consumer app behavioral analytics
Product analytics, user behavior
- Slack, Gmail
- 29 GitHub repos
- Stripe, Postgres
- Drive, Figma, Linear
Classes present
Company operations corpus
Full company corpus
- Code repositories
Classes present
Startup codebases
Codebases
- Credit agreements
- Offering documents
- Fund formation
- Operational
Classes present
Private-markets legal documents
Legal work product
- ~10,000 profiles
- Food journals, images
- Meal, workout, Apple Health
- In-app chat
Classes present
Consumer health and nutrition logs
User health and nutrition data
- BMR / BPR
- QA release
- Deviations, CAPAs, OOS
- Regulatory correspondence
Classes present
Pharmaceutical batch records
Manufacturing records
- Generated alpha signals, 1.6 TB
- Other proprietary signals and research, 5.4 TB
- Strategy and cloud application logs, 3.1 TB
- Warehouse, op DBs, streaming, 59 repos, 249k Slack messages, Notion, Linear, Figma, 700 GB
Classes, to scale
Trading operations corpus
Full operational corpus