• Joined on 2026-07-18
CPU rules+cheap-ML cascade for job-title validity, with an independent post-build audit
Updated 2026-09-03 12:17:06 +00:00
Self-contained technical handbook for the job-title/contact-validation research line. Clone and feed to a new session; do not expect sibling-repo access.
Updated 2026-08-28 09:30:51 +00:00
Session handoff: everything built, evaluated, and published from this lab — context for starting a new coding session.
Updated 2026-08-26 17:42:02 +00:00
Deep research context: every job-title/vendor-contact validation project on this box — methods, measured numbers, retractions, open problems. Bootstrap doc for a fresh session on this specific problem.
Updated 2026-08-26 17:42:02 +00:00
Public subset of the knowledge base — auto-generated, do not edit directly
Updated 2026-08-23 01:53:50 +00:00
Job-title validation (not classification) for contact lists: DROP/USE/REVIEW with logged reasons. 0% false accepts, 0.7% false drops on a holdout eval. Local-only.
Updated 2026-08-21 18:12:05 +00:00
Master index of every project built/researched on this VPS — bootstrap context for a fresh Claude session
Updated 2026-08-21 06:26:44 +00:00
Phase 2: is a job title on a contact row valid or invalid, using company/country/name/email/phone context. Builds on job-title-validation, titlevalidate, titlebert, csuite_title_ontology, multilingual-job-title-matching.
Updated 2026-08-20 03:26:37 +00:00
Jupyter notebook: is this job title plausible AT THIS COMPANY? (L2 — legal form, size, cardinality; PLAUSIBLE is confidence-capped without a registry)
Updated 2026-08-19 23:35:32 +00:00
Jupyter job-title validation workbench: multi-signal input, VALID/INVALID + confidence, ESCO/O*NET/UK SOC/India NCO neighbours, looping practice sequences.
Updated 2026-08-19 17:50:01 +00:00
Multilingual job title matching: ESCO+O*NET data, lexical baseline vs frozen multilingual embedding matcher, honest held-out eval with real negatives.
Updated 2026-08-19 12:52:12 +00:00
Validates 3rd-party-vendor contact/profile data (JSON + tabular, public/private companies) for job-title plausibility and identity-level fake-profile risk. Tiered confidence output, never a bare true/false. 10-step build, every claim measured against real registries or public data.
Updated 2026-08-17 02:23:39 +00:00
CPU-reproducible ML tutorials: LoRA fine-tuning, mmBERT embeddings for job-title/profanity detection, TalentCLEF/JobBERT-V3 survey, ESCO/O*NET model research. Every number measured on a 4vCPU/8GB VPS, no GPU, no API keys.
Updated 2026-08-16 00:01:36 +00:00
mmBERT-small multi-head encoder classifier for job-title validation (validity/type/toxicity/language), Grok-CLI teacher distillation, CPU-only training+ONNX serving
Updated 2026-08-15 04:55:25 +00:00
TitleBERT for Apple M5 Pro (48 GB): multilingual job-title classifier, bundled data, pass/fail verify gates, production file scorer. Public clone-and-run package.
Updated 2026-08-15 03:43:19 +00:00