πŸŽ‰ PrimeSkills LiveVisit β†’
⚑Solo Founder · Independent Creator

Hi, I'mDongqi Chen
Building Interesting Things Here

10+ years in automotive systems, AI application explorer, indie developer. Sharing hands-on experience in in-vehicle development, AI engineering, and team management, while incubating small but beautiful products.

CarOSAI/LLMIndie Hacker
BLOG

Latest Articles

Technical articles, tutorials, and insights covering automotive dev, AI engineering, and team management

LLMAutomated TestingAI Engineering

LLM-Driven Automated Testing Series (14): Production Engineering β€” Why Testing Turns Out to Be the Cheapest Line Item in This Whole Series

A code-generation workflow I documented earlier burned thousands of dollars in tokens during debugging. The testing scenarios verified across this series β€” AutoRestTest testing a 15-operation API for about a dime, pytest-triage capping itself at 10 model calls per run by default β€” are off by orders of magnitude. This article puts the evidence from 13 prior articles side by side to answer the three questions the planning doc left for last: why testing-scenario cost is what it is, where the human-in-the-loop gate actually belongs, and which scenarios should explicitly avoid LLMs altogether.

Β·12 min
LLMAutomated TestingFlaky Test

LLM-Driven Automated Testing Series (13): Flaky Tests Aren't a Rerun-Count Problem, They're a Classification Problem

pytest-rerunfailures (477 stars) and flaky (397 stars) handle test instability with the same one-liner: if it fails, rerun it a few times. That doesn't answer what the planning doc actually asks β€” was this failure a real bug, environment jitter, or a badly designed test? The only project actually attempting that root-cause classification with an LLM is a 2-star, less-than-two-months-old repo called pytest-triage β€” and its most instructive feature turns out to be the four safety invariants it wraps around AI judgment, guaranteeing AI is never allowed to affect whether a test actually passes or fails.

Β·11 min
LLMAutomated TestingVisual Regression

LLM-Driven Automated Testing Series (12): Beyond Pixel Diffs β€” Who's Actually Making LLMs 'Understand' UI Changes

BackstopJS, at 7.2k stars, handles rendering noise with exactly four numeric config knobs β€” misMatchThreshold, requireSameDimensions, ignoreAntialiasing, usePreciseMatching β€” which is still just fighting font-rendering jitter with pixel-level tolerance. The project that actually wires a VLM/LLM into the visual-regression judgment loop is a 23-star repo called vlmkit, whose own internal benchmark notes state outright that one model hallucinates red into red. Putting these two side by side answers, concretely, what semantic visual regression actually solves β€” and what new problems it introduces.

Β·12 min
LLMAutomated TestingAPI Testing

LLM-Driven Automated Testing Series (11): Why API Testing Is Actually the Most Conservative Battleground for LLMs

Keploy has 18.5k stars and its README says 'Expand API Coverage using AI' β€” but grep the entire open-source Go codebase for LLM/AI keywords and you'll find zero implementation code. Schemathesis, at 3.6k stars, uses no LLM at all β€” it finds real server 500 errors purely through Hypothesis-based property-based fuzzing. The projects that actually call an LLM API in their core logic turn out to be two sub-100-star projects: AutoRestTest and api-automation-agent. This article pins down why: structured API input/output is naturally suited to traditional fuzzing, so what can LLMs actually do here that traditional tools can't β€” and why does this domain stay cautious instead of going all-in the way UI automation has?

Β·13 min
LLMAutomated TestingSelf-Healing

LLM-Driven Automated Testing Series (10): Self-Healing Locators β€” Where Heuristic Scoring Hits the Ceiling of Semantic Understanding

Healenium is the most established open-source project in self-healing locators, and its README claims it 'leverages machine learning' β€” but reading the code directly reveals the core mechanism is a weighted scoring formula over DOM tree similarity, with hardcoded constants for weights and thresholds, nothing to do with machine learning. More notably, its 2026 addition of an AI-based XPath generation endpoint is gated behind a hard-coded exception: 'you must have a paid hlm-ai service.' This article first cracks open exactly how Healenium's scoring formula works, then examines how MarketSquare/robotframework-selfhealing-agents β€” a genuinely LLM-based self-healing project β€” designs its multi-agent architecture, and finally answers the question the planning doc raised: where exactly does the tradeoff between healing success rate and false-positive rate actually bite.

Β·14 min
LLMAutomated TestingMetaGPT

LLM-Driven Automated Testing Series (09): Mobile Automation (V) β€” When a Generic Agent Framework Sinks Down into a Testing Scenario

MetaGPT is a 'multi-agent software company' framework built for requirements analysis and code generation, but it ships a built-in Android testing demo. Reading the source directly overturns one premise from the planning doc: this demo does not reuse MetaGPT's usual Product Manager/Engineer/QA role division β€” it has exactly one custom Role, reusing instead the lower-level Team/Environment/Role/Action scheduling scaffolding. More importantly, the README states plainly that it 'referenced ideas and code from AppAgent' β€” this piece breaks down exactly what this generic framework inherited from AppAgent, what reusing generic scaffolding actually bought it, and what it cost.

Β·15 min
PRODUCTS

Things I've Built

Independent products, from idea to launch β€” all by myself

PODCAST

Let's Talk

Sharing tech insights and life thoughts on podcasts

CONNECT

Find Me

Follow me on these platforms for the latest updates

βœ‰οΈ

Subscribe to Newsletter

Curated content delivered to your inbox weekly β€” tech insights, product updates, and industry trends. No spam, unsubscribe anytime.

We only collect your email for the newsletter and will never share it with third parties. Unsubscribe anytime.

1,024 subscribers Β· Privacy-first, never shared