memelab BUILD NOTES

PRODUCT + IMPLEMENTATION

How one comment
becomes the best options.

The live product compares meaning, social dynamics, and tone—not just words. It retrieves 30 likely matches from 1,000 memes, ranks the best three, then explains each choice.

Try a comment
Live now: 1,000 reactionsSemantic retrieval · three scored options · serious-content guard · feedback
PERSONAL TOOL

THE LIVE 1,000

Retrieve. Classify. Explain.

Each reference has a name, meaning, relational pattern, example, tags, provenance, and media status. Vector search narrows 1,000 records to 30 before a dedicated classifier ranks the best three.

  1. 01Validate

    Trim and bound the comment. Treat it as untrusted text.

  2. 02Retrieve 30

    Use BGE embeddings and Vectorize to find the closest semantic matches.

  3. 03Guard and rank

    Jev checks sensitive help-seeking text and ranks the three closest social reactions.

  4. 04Explain three

    Llama writes specific explanations and fit scores without changing the ranked IDs.

THE RETRIEVAL SYSTEM

The path scales to 30,000.

At 3,000, retrieval keeps the prompt bounded. At 30,000, a first routing decision chooses meme, movie dialogue, or no reaction before searching the matching corpus.

Workers AI

Meaning, not keywords

The same embedding model converts catalogue records and incoming comments into comparable vectors.

Vectorize

A bounded shortlist

The vector index returns 30 likely candidates from hundreds or thousands of records.

Jev classifier

Tone and relationship

A fast classifier checks whether humour belongs and ranks candidates by social fit rather than keyword overlap.

D1 feedback

Evidence, not vibes

Yes/no feedback is retained for evaluation; it does not silently retrain the live model.

THE EXPANSION PLAN

30 → 300 → 1,000 → 3,000 → 30,000

300 references are live in the personal tool. Evaluation now guides metadata cleanup and the next expansion.

1,000 / 3,000 live“I love you 3,000” is the north star.
2,700 memes remaining · 10% complete
  1. 30
    Seed control

    Curated metadata, direct all-candidate comparison, live public feedback.

    LIVE
  2. 300
    Common reactions

    Semantic retrieval, scored multi-option results, searchable collection, and feedback.

    LIVE
  3. 1K
    Broader vocabulary

    Seven hundred new records add long-tail reactions with searchable previews and meaning-specific metadata.

    LIVE
  4. 3K
    I love you 3,000

    The next 2,000 will expand the long tail while the same retrieval and safety gates protect relevance.

    BUILDING
  5. 30K
    Movie dialogue reactions

    Route to a separate dialogue corpus first, then retrieve and rank within it so memes do not crowd out the right line.

    NEXT CORPUS

THE EVAL DATASET

110 held-out situations guide the next jump.

The original and expansion sets cover 50 situations. The 1,000-meme set adds 60 cases: 45 humour and 15 no-meme prompts, with labels independently audited twice and still pending owner review.

Latest draft: the 1,000-meme path reached 96% retrieval, 87% top-one relevance, 96% top-three relevance, and 93% correct serious-content abstention. Labels remain pending owner review.

THE EXPANSION RULE

Bigger is useful only if the results stay good.

The personal catalogue stays available while evaluation tracks top-three relevance, correct no-meme decisions, valid IDs, and latency. Comments are processed by Cloudflare Workers AI and classifier.dev; previews load from their listed source hosts.