Skip to content

Add apify.lazada.search.products (Lazada SEA product search) - #641

Closed
herus13 wants to merge 1 commit into
superdesigndev:mainfrom
herus13:vendor/lazada
Closed

herus13 wants to merge 1 commit into
superdesigndev:mainfrom
herus13:vendor/lazada

Conversation

@herus13

@herus13 herus13 commented Sep 23, 2026

Copy link
Copy Markdown

Lists our public Apify actor herus13~lazada-scraper as a catalog endpoint under the existing apify provider. Search Lazada across six Southeast Asian sites (SG, MY, VN, TH, PH, ID); each row carries title, price, price_original, currency, rating, review_count, sold_count, seller_name, location, brand, product_url.

Contact: bootforge.ai@gmail.com

Files

  • src/treg/catalog/apify.yaml — new endpoint apify.lazada.search.products + proposed_capabilities: lazada.search.products
  • src/treg/catalog/capabilities.yaml — new lazada platform (E-commerce); the taxonomy had no Lazada entry
  • src/treg/catalog/examples/apify.lazada.search.products.json — one real, scrubbed product

Verification evidence ledger

endpoint http test target metered catalog price matches? evidence
apify.lazada.search.products 201 search_terms:["iphone 15 case"], region sg, maxItems=1 usageTotalUsd 0.00588 (compute); PPE self-unbilled per_result $0.005 ⚠️ hit price unobserved run FoMZQgBrce1dYgbik, 1 product returned; price read live from actor pricingInfos (eventPriceUsd 0.005, sole event, apifyMarginPercentage 0.20)
  • Bad-key probe (apify provider probe, GET /users/me): garbage token → HTTP 401, body {"error":{"type":"user-or-token-not-found","message":"User was not found or authentication token is not valid"}} (2026-09-23).
  • scripts/catalog_validate.py: OK — 0 error(s), 0 warning(s).
  • Tests: test_catalog_validate + test_additional_capabilities + test_catalog_drift = 121 passed locally (Python 3.13). Full-suite server/DB-fixture tests not run in my environment; CI covers them.
  • No credential value anywhere in the diff.

Cost provenance note (honest ⚠️)

This is a pay-per-event actor: one product-result event = one product = $0.005, read live from the actor's pricingInfos. Cost is stamped confidence: documented / source: rate_card_api, not verified, because I am the actor owner: Apify does not bill an owner the per-result rental on their own actor, so chargedEventCounts came back null and the caller-side charge is not observable from my meter. A reviewer with any other Apify token will meter the $0.005 on a real run. Please re-verify and, if the observed charge differs, correct the YAML to the metered value (source: observed).

Full-surface map

The actor exposes exactly one operation — a keyword product search (run-sync-get-dataset-items), catalogued here. Async large pulls are already served generically by the provider's apify.web.scrape.job.start/status/results. Nothing excluded.

Billing model

BYOK: the caller runs the actor on their own Apify token; pricing is Apify's standard pay-per-event ($0.005/product, 20% Apify margin). Actor page: https://apify.com/herus13/lazada-scraper — pricing: https://apify.com/pricing

List the herus13~lazada-scraper Apify actor as a catalog endpoint: search
Lazada across six Southeast Asian sites (SG, MY, VN, TH, PH, ID), returning
price, rating, review/sold counts, seller and location per product.

Adds the `lazada` E-commerce platform to capabilities.yaml (the taxonomy had
no Lazada entry) and maps the endpoint to a new `lazada.search.products`
proposed capability.

Verified 2026-09-23: run-sync capped at maxItems=1 returned HTTP 201 with one
real product (run FoMZQgBrce1dYgbik). The $0.005/product price is the live
Console PPE rate (pricingInfos eventPriceUsd 0.005, sole event); the owner
meter cannot observe the caller-side charge, so cost stays confidence:
documented. catalog_validate.py exits 0.
@JayZeeDesign

Copy link
Copy Markdown
Contributor

Thanks for this listing. It's included, with your co-authorship, in #657, which combines your four actor PRs and adds the billing code needed to serve them on treg's shared key (runs are capped by maxTotalChargeUsd and settled by returned rows plus the per-run start/compute charge). Closing this one in favour of #657.

JayZeeDesign added a commit that referenced this pull request Sep 26, 2026
…pped billing (#657)

* fix(catalog): keep Apify actor starts off treg's shared key

apify.web.scrape.job.start is priced free because the run bills later by the
actor's own pricing, and nothing on the shared-key path meters that run. On
treg's key it let any caller run any actor on treg's Apify account at no
charge. It now needs the team's own Apify key; job.status and job.results
stay open because per-team ownership already scopes them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): add Weibo, TikTok Shop, Lazada and Google Maps Apify actors

Combines #641-#644 into one catalog change: nine run-sync entries over four
herus13 actors, plus the lazada platform. The six Weibo entries serve on
treg's key at a price observed on it. Lazada, Google Maps and TikTok Shop are
own-key only: their actors bill run compute or a per-GB start fee that a
per-result price cannot meter.

Co-authored-by: Herus13 <bootforge.ai@gmail.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(catalog): require own keys for Weibo actor runs

* test(frontend): isolate landing interactions from continuous WebGL rendering

* feat(call): let platform_request pin query parameters

Some upstreams take their spend bounds as query options rather than body
fields (Apify's maxTotalChargeUsd, memory and timeout run options). A
queryParams pin must be sent exactly once and is compared as the pinned
value's type; own credentials keep the upstream contract.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): bill Apify platform calls by returned dataset rows

run-sync-get-dataset-items answers a bare array, so every Apify per_result
call settled at its estimate whatever it returned. Count the rows, add an
optional per-row-independent call_fee for the actor's start or compute
charge, and require maxItems (1-200) on the platform key so the hold is the
worst case. Apify's usageTotalUsd lags a finished run by minutes, so the
response body is the settlement evidence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): serve the herus13 Apify actors on treg's key

Weibo, Google Maps, TikTok Shop and Lazada now settle on the platform key by
counted dataset rows plus a flat call_fee (actor start, or Lazada's run
compute). memory and timeout are pinned so the fee is fixed and a run ends
before Apify's synchronous wait; TikTok Shop is held to keyword search.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): bound Apify platform calls by maxTotalChargeUsd

maxItems does not bind pay-per-event actors whose own input sets the row
count, so a one-row hold could settle thousands of rows. Require the
maxTotalChargeUsd run option Apify enforces (at most $1), hold it plus
call_fee, accept only the run options each once in plain ASCII, and bill the
hold when a run reaches its cap. Pin meta-ads enrichment off and bill the
LinkedIn actor-start event on treg's key.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): cap herus13 Apify runs by maxTotalChargeUsd on treg's key

Row notes name maxTotalChargeUsd as the enforced spend limit; Lazada's
compute fee is 0.015 under a 180-second timeout pin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): name maxTotalChargeUsd as the Apify spend cap

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): treat an Apify run within two rows of its hold as capped

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(call): require a short timeout and a three-row cap on Apify platform runs

Past Apify's 300-second synchronous wait a run answers 408 and keeps billing,
so every Apify per_result platform call now names timeout <= 280. A cap under
call_fee plus three rows would bill an empty answer in full, so it is refused.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(call): keep Apify platform runs inside every wait

treg's upstream read timeout (call_timeout_s, 180 s) and the MCP client's
120 s end the call before a 280 s run finishes, releasing the hold unbilled
while the run keeps billing. Bound timeout to 90 s and 30 s under
call_timeout_s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): pin herus13 Apify runs to a 90-second timeout on treg's key

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): bill a timed-out Apify platform run at its hold

A run that outlives its own timeout answers 400 run-failed with no rows, but
Apify billed its events up to the caller's cap and the run id in that body
reads the dataset. The caller chose the run's size and timeout, so the hold
settles instead of releasing. The minimum-cap check compares micro-USD.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): tell Lazada callers to keep platform runs small

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): release timed-out Apify runs; the account must stay Restricted

Billing the whole cap on a TIMED-OUT 400 overcharged callers: a run that
timed out after its start event cost Apify $0.00005 and would have billed the
full cap. The loss treg absorbs stays bounded by the $1 cap, and keeping the
Apify account's resource access Restricted stops anyone reading the unbilled
run's rows by id. Examples now show the 90-second platform timeout.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(money): name disconnects as a bounded Apify loss; call_fee wording

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): record Lazada's 0.05 minimum cap and 10-product floor

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(asynctasks): keep the Bright Data and CompanyEnrich platform keys the merge dropped

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(catalog): verify every Apify per-result price on the platform key

Each row's test_request ran on 2026-09-26 and Apify's chargedEventCounts
matched its rate card. The TikTok ad library actor returns rows again, so it
is verified with a captured example.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): route Google Maps and Weibo post detail through Apify

Adapters let treg.google.serp.maps and treg.weibo.post.detail choose the
Apify actors, with the run's spend cap, timeout and memory fixed so the
platform guard accepts the child call.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(catalog): drop Lazada's compute fee and keep account details out of notes

The actor's 2026-09-25 pricing no longer bills run compute to the caller (a
79-second run showed no platform usage), so its call_fee over-charged every
call. Notes describe observations by price tier, not by the account that
made them, and carry no run ids.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): state tiered Apify observations without the account's tier

Notes for plan-tiered actors record the events billed and that they matched
the rate card for the key's tier, not the dollar figure that would name it;
the repeated Weibo observation and the owner-meter notes go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): bill LinkedIn's actor-start once per query it runs

The LinkedIn jobs actor bills its actor-start event for every job title x
location searched, so a flat call_fee under-billed any multi-query call.
cost.call_fee_per names the body arrays whose lengths multiply the fee. The
apify.yaml header no longer names a plan or calls the TikTok actor broken.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(money): describe call_fee_per

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(catalog): accept only declared fields on LinkedIn job search

The actor bills an actor-start for every geo id it searches, and geoIds was
undeclared, so it passed through unbilled. The row now takes its declared
filters only (salary, easyApply, under10Applicants and industryIds added);
places go in locations, which call_fee_per counts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): leave headroom above the Apify spend cap

A run whose charges land exactly on maxTotalChargeUsd can be aborted by
Apify and answer 400 with no rows, which releases the hold. Lazada's
10-product floor costs exactly its 0.05 minimum cap, so its example now uses
0.06 and the notes say to set the cap above the expected spend.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Herus13 <bootforge.ai@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants