Sofue Data

Independent data studio · Japan

Public data is overwritten every day. We keep the history.

Sofue Data takes daily snapshots of public, non-personal data that no one archives, such as app store rankings and company hiring activity, and turns them into pay-per-result data products. The pipeline is designed, built and maintained with Claude.

US App Store · Top Free · daily rank
PineDrama – Short Dramas
Aug 8 – Oct 6, 2026
#1 #25 #50 #70 Aug 8 Sep 1 Oct 1 #62 #7
Apple's public feed shows only today's chart. This line exists because we recorded every snapshot ourselves.
36
daily snapshots since Aug 8, 2026
76,000+
rows recorded
4,268
distinct apps seen across 10 storefronts
9,100+
open roles at 53 companies in the hiring registry

As of October 7, 2026.

Products

Clean data, priced per result

Each product starts from an official public endpoint, adds the history or normalization that the source does not provide, and charges only for the rows you receive.

Live on Apify

App Store Chart History

Top Free and Top Paid, top 100, for 10 storefronts: US, JP, GB, DE, FR, KR, CA, AU, BR and IN. Pull the current chart on demand, or the accumulated daily history for trend and momentum analysis.

Pricing
$1 per 1,000 results
For
Indie developers, ASO tools, investors tracking category momentum
Live on Apify

ATS Hiring Signals

Job postings from Greenhouse, Lever, Ashby, SmartRecruiters and Recruitee in one schema. Delta mode returns only newly opened and closed roles with first-seen dates, plus per-company rollups that show who just started hiring, and for what.

Pricing
Pay per event
For
Recruiters, B2B sales teams using hiring spikes as buying signals
Planned

Agent-payable Data API

The same history and derived signals, served over HTTP to AI agents that pay per request with x402, without accounts or API keys.

Pricing
Per request
For
Autonomous agents and developer tools

How it works

Capture today, sell the history tomorrow

Capture

A scheduled job reads each source's official public endpoint once a day. Today that means 20 App Store chart feeds and the Hacker News front page.

Accumulate

Rows are appended as NDJSON, one file per source per day, and committed to git so that every snapshot carries a verifiable timestamp.

Deliver

Apify actors sell current data and history per result. Derived signals such as rank percentiles, momentum and hiring deltas are built on the same store.

Parsing code can be copied in an afternoon. A day of history that was never recorded cannot be bought back.

One row from the Oct 6, 2026 snapshot
{ "ts": "2026-10-06T21:00:02.048Z", "country": "us", "chart": "top-free",
  "rank": 7, "appId": "6754134922", "name": "PineDrama - Short Dramas",
  "artist": "TikTok Ltd.", "releaseDate": "2025-12-23",
  "url": "https://apps.apple.com/us/app/pinedrama-short-dramas/id6754134922" }

Built with Claude

One operator, a team of Claude agents

The studio runs on Claude Code. Three specialized agents do the scouting, building and repair work that would otherwise take a team, which keeps the portfolio maintainable by one person.

Niche Scout

finds sources

Searches for long-tail public data, opens every candidate URL to confirm it is reachable, checks robots.txt and terms, and measures how crowded the marketplace already is.

Actor Scaffolder

builds products

Turns a target spec into a deployable Apify actor: typed input schema, pay-per-event billing, timeouts and retries, Dockerfile and README.

Maintenance Monitor

keeps them running

Reads run logs, separates transient errors from structural breaks such as schema drift, and proposes the exact parser patch, ranked by what is at risk.

Data principles

Boring, public, and clean

Non-personal data only

No names, contact details or anything that identifies a person.

No gated sources

Nothing behind a login, a paid license, or terms that prohibit collection.

Official endpoints first

Public JSON APIs and open data before HTML pages, with robots.txt respected.

Verifiable provenance

Every snapshot is timestamped and kept append-only.

What's next

Roadmap

  • More sources with no public archive. Transit and public-infrastructure open data are under evaluation.
  • Derived signals on top of the history: rank percentiles, momentum and category trends.
  • A self-serve history API and the agent-payable endpoint.
  • A self-healing maintenance loop across the whole actor portfolio.

Need a dataset that nobody keeps?

Questions, data requests and partnerships are welcome.

hello@sofuetakuma.uk