Tell us what you need. We'll collect, enrich, and deliver your private dataset — with a 100-row free sample and permanent archive access.
Social Intel collects and enriches public social media data from 100+ platforms on demand. Describe the data you need — topic, keywords, date range, and field requirements — and our platforms are auto-selected from your chosen category (and you can trim up to 20%). Our pipeline handles collection, normalization, sentiment analysis, topic classification, entity extraction, and quality scoring. Every dataset is enriched with the fields relevant to its platforms — drawn from 250+ available — and is delivered in CSV, JSON, JSONL, and Parquet formats. Pro subscribers get custom datasets at no extra charge.
Request for review
We'll collect your dataset and notify you via email. The full download requires an individual purchase (from $18), with permanent access. 100-row free sample always included.
Free
Build progress20%
Complete the form to build your dataset — 80% to go.
Your dataset
NameUntitled dataset
Categorygeneral
Fields15 configured
Rows10K–25K rows · $18
FormatsCSV · JSON · JSONL · Parquet
Build progress20%
80% left — finish your spec.
Pricing
Selected tier: $18 — pay once, permanent access. Pro plan — 7-day free trial — includes free custom datasets.
Enrichment Pipeline
Every custom dataset goes through our deterministic enrichment pipeline, which draws on 250+ available fields and applies the ones relevant to its platforms.
1. Auto-Select Platforms
Platforms are auto-selected from your chosen category across 100+ supported sources — you can review and trim up to 20% before submitting.
2. Generate Dataset Description
AI writes a concise 1-2 sentence description of the dataset based on your input, category, and selected fields. Used for catalog display and metadata.
3. Expand Search Queries
AI generates optimized search queries that scale with your selected row tier — from ~65 at the 10K–25K tier up to 400+ at 75K–100K — covering multiple angles: tutorials, comparisons, opinions, news, how-tos, best-of lists, and trends. For large counts, queries are generated in parallel batches to guarantee full coverage without truncation.
4. Optimize Field Selection
15 CORE fields (id, content, url, author, engagement metrics, etc.) are always included for data quality and consistency. Additional field groups auto-suggest based on description keywords — sentiment, financial, technical, health, gaming, and more.
Enrichment Pipeline
Smart Scheduling: Collection tasks are prioritised and queued. Each platform is scraped using custom-built adapters.
Post-Level Enrichment: Each collected post is analysed for sentiment, emotion, topics, financial signals, and domain-specific features (crypto tickers, game titles, health conditions, etc.).
Deduplication: Content is deduplicated by hash across platforms to ensure clean, unique data.
Multi-Format Export: Delivered as CSV, JSON, JSONL, and Parquet with schema documentation.
Deterministic enrichment — every field reproducible from source code.