Maison KOSÉ listing-page scraper (Playwright, Selenium, Puppeteer, or the 2Captcha Scraping Browser API via CDP) — product listings, prices, brands, stock and Japan's two consumption-tax rates, captcha solving, proxies, fingerprints.
kose.com resolves for nobody. Measured 2026-09-18:
$ dig +short @8.8.8.8 kose.com A # SERVFAIL, no answer
$ dig +norecurse @a.gtld-servers.net kose.com NS
kose.com. 172800 IN NS dns1.supremecenter.com.
kose.com. 172800 IN NS dns2.supremecenter.com.
$ dig +norecurse @162.210.102.140 kose.com SOA # REFUSED
The domain is registered (GoDaddy, created 1998-11-30, paid to 2029) and delegated to two nameservers that no longer serve the zone, so every resolver returns SERVFAIL. That is a broken delegation, not a bot manager: there is nothing at that name for anyone to fetch.
Kosé's actual online store is maison.kose.co.jp
(Maison KOSÉ) — which is where www.kose.co.jp sends a visitor, with its own
<meta http-equiv="refresh">. That is what this scraper reads.
Stated plainly rather than sold. Measured 2026-09-18 from a datacentre
address (netcup, Nuremberg, AS197540), against
/site/cosmedecorte/c/c15_p3/:
| request | result |
|---|---|
curl/8.5.0 |
HTTP 200, 90,310 bytes |
python-requests/2.31.0 |
HTTP 200, 90,310 bytes — byte-identical |
| a desktop Chrome User-Agent | HTTP 200, 90,310 bytes — byte-identical |
no User-Agent header at all |
HTTP 200, 90,310 bytes — byte-identical |
| 12 pages back to back, no delay | 12 × HTTP 200 |
and zero captcha markers of any vendor across 13 captures — no reCAPTCHA,
hCaptcha, Turnstile, DataDome, PerimeterX, Incapsula, Kasada or AWS WAF, no
data-sitekey, no challenge iframe, on listings, product pages, the site's
own 404 and an exhausted facet alike.
The grid is server-rendered into the first response, so a plain HTTP client gets the whole page. The three browser engines are a convenience here, not a necessity.
What the paid products would buy you, if that changes or at volume:
- 2Captcha solving — if Kosé ever renders a challenge. Nothing in this
repo says a captcha here could not be solved; there is simply none to
solve today. The solver is wired and the run pays for it under the usual
--solve-captcha when-blockedrule. - Proxies — many addresses, for a run large enough that one would be conspicuous. A Japanese exit is not required: a German datacentre address is served the same pages.
- The Scraping Browser API — a managed browser with a persistent profile and no local Chromium to maintain.
- Fingerprints — a consistent device identity. Verified working here (a JP fingerprint applied, 24 products returned), not needed.
git clone https://github.com/2scraper/kose-scraper
cd kose-scraper
python3 -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt # core: beautifulsoup4, requests
pip install -r requirements-playwright.txt # then ONE engine
playwright install chromiumInstall exactly one engine. playwright and pyppeteer pin incompatible
pyee versions and pyppeteer and selenium collide on urllib3; they do run
side by side in practice, but pip check reports the conflict and pip may
resolve it by downgrading something you wanted. Use a virtualenv per engine
if you need more than one.
# a brand's catalogue, three pages
python playwright_scraper.py \
--url "https://maison.kose.co.jp/site/cosmedecorte/c/c15/" --pages 3
# the same thing by category id — the id is the only thing that selects one
python playwright_scraper.py --category c15 --pages 3
# a tag facet, which crosses brands
python playwright_scraper.py --tags "シワ改善,スキンケア" --pages 2
# one product page: adds volume, subcategory, colour count and release date
python playwright_scraper.py --mode product \
--url "https://maison.kose.co.jp/site/onebykose/g/gMUSC/"
# no local browser at all
python scraper_api_client.py --key "$TWOCAPTCHA_KEY" \
--url "https://maison.kose.co.jp/site/cosmedecorte/c/c15/"Output is kose_products.json + .csv plus a kose_products.meta.json
sidecar. python3 env_config.py prints what was picked up from .env,
without printing secrets.
Measured 2026-09-18 across four live runs — 216 rows over the category route, the tag route and a product page:
| column | coverage |
|---|---|
title url sku brand |
216 / 216 |
price currency tax_rate |
216 / 216 |
in_stock image_url |
216 / 216 |
Prices ran ¥495 to ¥264,000. A 3-page run of /c/c15/ returned 72
products, 72 distinct skus, status complete, exit 0.
All three engines agree. Playwright, Selenium and pyppeteer were run against the same two pages on the same day: 48 of 48 rows identical on sku, price, brand and stock.
tax_rate— Japan runs two consumption-tax rates and the prices here are tax-INCLUSIVE. 10% is standard; 8% is the reduced rate on ingestibles, so a beauty supplement sits at 8% while the cream beside it is at 10%. Two tiles in 171 carried it. Computing a pre-tax figure from the wrong rate is wrong by two percent, silently, which is why it is a column.in_stock— read from the site's own button text as an ALLOWLIST, so a wording Kosé adds tomorrow reads as unknown rather than as a quiet "in stock". Verified in both directions: 19 of 171 measured tiles say 現在購入頂けません or 予約受付終了.price_is_store_price— the site marks its own selling price with※, expanded on a product page asMaison KOSÉ販売価格. It says which of two different things a number is.variant_of— on a product row, the sku the page's ownrel=canonicalpoints at when that is a DIFFERENT product. It groups two sizes or shades of one item, and it exists as a column precisely so nobody uses the canonical as the row's URL — see the traps below.badges— the site's ownnew/limited/onlineflags, on 21% of rows.limitedis often why a row vanishes between two runs.
A column null on every row of every run should not exist, and removing one needs the measurement written down:
- no
original_price/discount_pct/lowest_price_30d— 360 price nodes surveyed and not one carries a struck-through, was- or reference price. There is no discount chain on this site, so there is no second view to reconcile a price against. - no
rating/review_count— absent from every tile and every product page's JSON-LD. The site publishes staff opinion pieces, which are editorial content by named employees rather than a score of anything. - no
locale— one market. Nohreflangset, no/en/,/site/en/or/global/route (all 404), JPY on every row. - no
sort— the site offers no ordering control on either listing route, so there is nothing a run could have asked for.
Each of these cost time to find and each is guarded by a check.
The brand in a URL selects nothing. /site/xyz/c/c15/ returns
byte-identical bytes (97,017) to /site/cosmedecorte/c/c15/, and the page's
own rel=canonical rewrites it. Only the cNN addresses a category and only
the gSKU addresses a product. So brand is read from each TILE — and a tag
listing genuinely mixes brands, 11 of them across 72 rows on a live run.
A category listing walked past its end serves its LAST PAGE AGAIN. Not a
404 and not an empty grid: /c/c15_p14/ and /c/c15_p15/ each return the
same 71,199 bytes as /c/c15_p13/, HTTP 200, with 16 real tiles on them. A
run that stopped on "a page came back empty" would never stop. This scraper
plans against the site's own N/Mページ counter — /c/c15/ states 13 pages,
and 12 × 24 + 16 = 304 products is the whole of Cosme Decorté — and still
stops on "this page added no new sku". A TAG listing ends differently, with a
genuinely empty page, and publishes no counter at all.
The wrong pagination parameter fails silently. On a tag listing ?p=2
works while &pageno=2 and &page=2 are ignored and return page 1 with HTTP
200. A run that built one of those would dedupe every page down to page 1's
rows and report a complete run holding a single page.
A quarter of product pages canonicalise to a different product. Measured
over 20: /g/gJQBA/ is the 60 g cream and its rel=canonical points at
/g/gJQGA/, which is the 20 g one. Five of the twenty do this. A row built
from the canonical would link to the wrong product while its title, sku and
price described the right one, so the row's URL is built from the page's own
sku and the canonical goes in variant_of.
A product page reuses the listing's tile class. c-product__item appears
on <div>s for staff-review cards — 15 of them on one measured page — plus
an awoo recommendation block. A class-only selector would invent fifteen
products per product page, each with a /staff/ URL and no sku. The tile
selector is element-qualified (li.c-product__item) and sees zero.
Product-shaped links outnumber real tiles. A grid page carries 48–57
/g/g…/ hrefs against 24 tiles (and 107 under Selenium, where more of the
page's JavaScript has run). The extra links are carousel banners. Rows come
from tiles; the link count is only used to tell a broken parser apart from an
empty category.
A 404 and an empty page are different answers, and the site gives both.
A bogus goods code or category id (/site/x/g/gZZZZZZ/, /c/c99999/)
answers HTTP 200 with the site's full chrome and no tiles — that is an
answer about the CATALOGUE. A wrong PATH answers a real 404 with a bare
1,040-byte page carrying no site chrome at all, and the scraper stops on it
rather than retrying: retrying spends the budget on an address that will
never exist. The bare tag-facet URL (/site/itemtags/list.aspx with no
tags=) is the one this scraper accepts that really does 404.
Selenium cannot read an HTTP status — WebDriver exposes none — so the parser
also reads the status the site states in its own <title>404- …. That is
what makes all three engines agree instead of two being better informed than
the third.
A run pointed at page 10 reads page 10. --url .../c15_p10/ --pages 3
fetches 10, 11 and 12; the start page is honoured rather than normalised
back to 1. The page column is the run's own index, so the sidecar's
start_url is what says where a run began.
Zero is a real price and a real rank. Prices reach six figures — ¥264,000 for one cream — so a parser that mishandles the comma turns that into 264. Every fixture in the suite pins the value, not the coverage.
- Selenium cannot use an authenticated remote CDP endpoint. chromedriver's
debuggerAddresstakes a barehost:portwith nowhere to put a password. Playwright and pyppeteer authenticate on the WebSocket upgrade; use one of those with--cdp-endpoint. - Selenium's
--proxy-servercannot authenticate at all. Credentials are stripped and a warning printed rather than letting you believe otherwise. - pyppeteer is effectively unmaintained and its own README points at Playwright.
- Never set a proxy or a fingerprint over
--cdp-endpoint. The remote browser brings its own; the engines refuse the combination on purpose. --concurrencyis refused over--cdp-endpoint. A Scraping Browser profile allows one live connection, so workers collide. Use severalpids, one run each.
Honest about the gap, because §13 of this family's playbook says a claim is measured or absent:
| path | status, 2026-09-18 |
|---|---|
| Playwright, listing / tags / product | verified — 72, 72 and 1 rows, exit 0 |
| Selenium, listing | verified — 48 rows, identical to Playwright's |
| pyppeteer, listing | verified — 48 rows, identical to Playwright's |
Scraper API (scraper_api_client.py) |
verified — 24 rows, 90,835 bytes, $0.0005 |
--fingerprint (2Captcha Fingerprint API) |
verified — JP fingerprint applied, 24 rows |
| credential masking in errors | verified — a deliberately wrong endpoint's password appears 0 times in the output, including in Playwright's own five-line call log |
--cdp-endpoint (Scraping Browser) |
not verified here — the profile available to this machine had expired (401 deny_no_user). Those credentials live about a day; the code path is the family's and is exercised by its siblings. |
--solve-captcha |
never fired — no challenge has been observed on this site at all |
0 ok · 1 crash · 2 bad usage · 3 blocked · 4 zero products ·
5 remote API error · 6 partial
A run that finds nothing writes nothing — it will not replace last night's
good output with []. Pass --allow-empty to opt out.
python3 smoke_test.py # offline, no engine library required
python3 smoke_test.py -v # and say what each check looked at
pytest # the same suite, wrappedFixtures are cut verbatim from real captures and were verified to parse to identical values to their untrimmed originals before being committed. The one staff-review card in the product fixture has its author name and text replaced with placeholders: it is there because the parser must IGNORE it, and republishing a named employee's writing is a separate act from the site showing it on its own page.
MIT. This repo reads public pages: listings, prices and product descriptions. It does not log in, does not touch a cart or a wishlist, and does not collect personal data.