How We Work — From Raw Data to Final Results

Nothing is hidden. Every engagement runs the same five steps — in replication mode (published data, checked against the paper) or build mode (your own data, verified against the data itself). The pipeline itself — script, data, report — is the deliverable.

PhD in psycholinguistics (UNSW), BA in linguistics (CUHK) — pipelines built to work the way research actually happens.

Who runs the work

Every engagement is staffed by a structured team of specialist AI roles before any step begins. I direct the team, make the key judgement calls, and deliver personally — one accountable human at the top, and a dedicated QA layer that reviews every output before it reaches you. Prefer to own the machinery? Set up your own isolated team →

  • Patrick Chu — founderDirection, judgement calls & final deliveryOne accountable person
    • Personal AI assistantoperations · tracking · briefings · cross-checks
    • Project managementplans scope, sequencing & decisions
      • Domain specialistsstatistics · methodology · subject area
      • Developerbuilds analyses, documents & tools
      • QA / claim reviewchecks every output against the evidenceGates every release
Review before release

▲ Every output passes the QA review layer before it reaches you.

The five steps below describe what this team does, in order.

  1. 1

    Know the data — public or yours

    What happens: Replication: locate and verify the public deposit (OSF, PLOS, GitHub). Build: inventory your dataset — structure, codebook, provenance.

    Why it matters: A fixed starting point means stable results; on new data, this inventory becomes the project's codebook.

  2. 2

    Load and sanity-check

    What happens: Row and column counts, value ranges, and missingness appear before any analysis begins.

    Why it matters: The first reviewer question is whether the data matches what the paper claims — answered visibly, up front.

  3. 3

    Build or rebuild the analysis

    What happens: Replication: core models rebuilt from scratch in an open statistical stack. Build: analysis constructed from your design in the same stack.

    Why it matters: Independence is the point — we reconstruct logic rather than echo an output.

  4. 4

    Verify — against the paper, or against the data itself

    What happens: Replication: published statistics versus the re-run, item by item. Build: a verification battery — diagnostics, sensitivity checks, recovery of known effects.

    Why it matters: Where numbers match, you see it exactly; where they differ, the difference is flagged — never explained away.

  5. 5

    Report and hand over

    What happens: Comparison table or verification report, plain-language summary, plus reproducible script and data.

    Why it matters: The deliverable is yours to keep — your team can re-run everything after the engagement ends.

Every pipeline follows the same discipline — load → check → analyse → validate → report. What you see in the demo is exactly what you get on your own data. Your data gets its own verification battery instead of a published answer key.

Pricing

Fixed-fee, USD-first — one approximate HKD anchor on the trial line only. Every engagement is scoped before work begins.

TierFrom
Trialfrom US$1,000 (≈ HK$8,000) — free if not useful
AI team setup (own your team)from US$1,000 (≈ HK$8,000) — one-time; isolated server provided by you
Small deliverablefrom US$2,600 (≈ HK$20,000)
Grant-scale projectfrom US$5,100 (≈ HK$40,000) (milestone-based)
Research retainerfrom US$1,300/month (≈ HK$10,000/month)
Admin retainerfrom US$640/month (≈ HK$5,000/month)

Want to see the whole machine, including the code? The fastest route is the trial — every script, documented, run on your own published data. See the FAQ →

Admin agent and teaching modernization each have their own process — also in the FAQ.