You commented WEEKEND. Here is the build order, a prompt to start each one, and the thing that kills all five if you skip it.
The order is not arbitrary. Each build leaves behind the thing the next one needs. Do them out of order and you will build the same plumbing three times.
1. The competitor ripper
2 to 3 hours. Teaches: scraping APIs and local speech to text.
In: ten accounts in your niche. Out: their best performing posts of the week, downloaded, transcribed, ranked.
What it leaves behind: a working scrape and transcribe spine. Build 2 reuses it.
ROLE: You are building me a competitor content ripper I can run daily.
TASK: Given a list of accounts, return their best performing recent posts with transcripts.
STEPS: Ask which platform and which ten accounts. Pick a scraping route that does not use my own logged in session. Pull post metadata and engagement. Rank by engagement against that account's own median, not raw likes. Download the media for the top N. Transcribe locally. Write one row per post to a file I can sort.
RULES: Never scrape from my own authenticated session, use a third party scraper. Rank against the account's own baseline or big accounts drown out small ones. Transcription runs on my machine, nothing uploaded.
OUTPUT: The script, the run command, the cost per run, and what breaks first when an account is private.2. The signal watcher
3 to 5 hours. Teaches: change detection.
In: a sheet of 200 people you would want to work with. Out: every morning, who moved, and a drafted opener about the move.
The hard part is not the writing. It is storing yesterday's version of the world and diffing it against today's. Once you can do that you can watch anything: pricing pages, job ads, who quietly dropped off a customer logo wall.
What it leaves behind: a snapshot store. Build 5 reads from it.
ROLE: You are building me a daily change detector over a fixed list of people.
TASK: Tell me who changed since yesterday, and draft one opener per change.
STEPS: Take my sheet. Each morning, fetch the current state of each row and write it to a dated snapshot. Diff today against yesterday on the fields I name. For each real change, draft a three sentence opener referencing the change itself. Skip everything unchanged.
RULES: Store the snapshot before you diff, never diff live against live. Ignore cosmetic changes, formatting and casing. A change I cannot act on is not a signal, so let me set which fields count. Never send anything, only draft.
OUTPUT: The snapshot schema, the diff logic, the drafts, and a count of changes ignored as cosmetic.3. The reply classifier
4 to 5 hours. Teaches: classification, and the eval set.
In: your inbox. Out: the messages from actual humans.
Out of office notices, auto acknowledgements and bounces all look like interest in a dashboard, and most people are counting them.
The real lesson is the second half. You hand label 200 replies and build an eval set, so you can prove the thing works instead of feeling like it does. Almost nobody does this, and it is the whole difference between a demo and something you would trust with your pipeline.
What it leaves behind: labelled data. Build 4 learns the shape of that habit.
ROLE: You are building me a reply classifier plus the eval set that proves it works.
TASK: Separate real human replies from automated noise.
STEPS: Define the classes first: human reply, out of office, auto acknowledgement, bounce, unsubscribe. Have me hand label 200 real examples before you write any classifier. Hold back 50 as a test set I never train or tune on. Then build the classifier and report accuracy per class on the held back 50.
RULES: The eval set comes first. If you write the classifier before I have labelled, you have built a demo. Report per class accuracy, never one overall number, because the rare classes are the ones that cost me. Tell me plainly which class it is worst at.
OUTPUT: The classes, the labelling sheet, the classifier, and a per class accuracy table on the held back 50.4. The lead scorer
6 to 7 hours. Teaches: embeddings and vector similarity.
In: your closed won list. Out: every new lead ranked by how much it resembles the deals you actually won.
Same technique behind every chat with your documents product, pointed at your customers instead of your PDFs.
ROLE: You are building me a lead scorer from my own closed won history.
TASK: Rank new leads by resemblance to deals that closed, not to my ICP document.
STEPS: Take my closed won list and my closed lost list. Build an embedding per account from the fields I actually had at the time the deal opened, not fields I only learned later. Score new leads by similarity to won, minus similarity to lost. Show me the top 20 and the bottom 20.
RULES: Only use fields available at the top of the funnel, or the score is cheating and will not generalise. If I have fewer than 30 closed won, say so and tell me the score is not yet meaningful. Show me why each lead scored where it did, in one line.
OUTPUT: The scored list, the reason per lead, and an honest statement of whether my sample is big enough.5. The memory layer
A full day. Teaches: entity resolution.
In: everything. Out: before anything goes out, it checks whether you have already spoken to that person, on any channel, in either direction.
You learn entity resolution, matching the same human across systems that disagree about who they are. It is the least glamorous problem in software and it is absolutely everywhere.
I have watched an automation message someone four times about a thing that was already cancelled. One lookup stops that.
ROLE: You are building me a contact memory layer that sits in front of everything that sends.
TASK: Answer one question before any outbound: have we already spoken to this person, on any channel, in either direction.
STEPS: Pull the last N messages from each channel I use. Normalise identities: lowercase emails, strip plus addressing, convert phone numbers to E.164, fuzzy match names only as a last resort and flag it. Build one record per human with every identifier attached. Expose one function: last_contact(person) returning the channel, the direction and the timestamp.
RULES: Never merge two people on a name match alone, flag it for me instead. Direction matters as much as recency: they replied last is a different situation from we sent last. Return "unknown" rather than guessing, because a false negative here sends a duplicate.
OUTPUT: The identity normalisation rules, the merge log with every flagged case, and the one function everything else calls.The part that did not fit in the reel
Build 6 is the one nobody builds, and it makes the other five worthless.
An SMS follow up sequence of mine ran every five minutes for 24 days and sent nothing. 6,912 runs. Zero messages left the machine.
The cause was one line of plumbing. The job appends to a log file. That log file was owned by a different user than the one the job runs as. The redirect failed, so the shell killed the command before any code loaded. No error, no alert, nothing to notice.
Then I checked the rest and found the same disease wearing another face: a reply watcher polling every 30 minutes and returning an error on every single call, because the account behind it had lapsed. It ran perfectly. It could never find a reply.
From the outside both looked exactly like working jobs.
The expensive part was not the downtime. It was the coverage. I had stopped checking by hand because something was watching. The thing watching was dead.
Four checks that tell you which of your jobs are lying:
Compare who owns the log file with who runs the job. A mismatch means it dies at character one.
Read the log's last modified time against the schedule. A five minute job with a three week old log is a corpse.
Read the last 40 lines. Repeated errors mean it runs and achieves nothing. Byte identical lines mean it is looping on dead state.
If two jobs share a file, check which one writes as the wrong user. Mine did, and an approval message reposted 36 times before I caught it.
Whatever you build from the list above, build the thing that watches it before you trust it.
Honest notes
The times are my estimates for a competent builder who has done something adjacent. They are not measured. Yours will differ.
I have built and run versions of 1, 2 and 5. I have not shipped 3 or 4 as production systems, so take those two as a design I would follow, not a result I am reporting.
Anything in this list that sends needs two rules on top: check who spoke last before you write, and send on the recipient's clock, not yours.
Build them in that order. Each one feeds the next.
Reply and tell me which one you started with.