Remix.run Logo
▲ vslovik 7 hours ago

Location: Pisa, Italy (CET) Remote: Yes, remote only Willing to relocate: No Technologies: Python, LLM/agent orchestration, local embeddings, retrieval evaluation, AWS (Lambda/SQS/EventBridge), Terraform, Django, PostgreSQL, LightGBM/PyTorch, NLP/Transformers Résumé/CV: linkedin.com/in/vslovik Code: github.com/vslovik/fenix — local-embedding search whose relevance is actually measured: labelled control probes, blinded human ranking, precision@k. No API keys. Email: valeriya.slovikovskaya@gmail.com

Software architect, 15+ years in production systems, almost entirely startups and internal startups — fintech, e-commerce, pharma, publishing.

The work I get pulled into is the recurring startup problem: a service shipped fast under launch pressure, without adequate tests, that later has to be made reliable without being stopped. Incident response, re-architecture, and the release discipline that keeps it from happening again. Most recently that has meant a regulated UK consumer-credit platform — loan servicing, arrears, forbearance, statutory breathing space, and early-settlement calculations written against consumer-credit legislation. Regulation as code, behind a test suite larger than the production codebase.

I've done that in all three configurations: taking a core system from problem statement to release, leading the team that carried it (1 to 7 engineers in ten months), and now doing the same work again with agentic tooling covering what the team used to.

On the data side: a LightGBM acquisition model over a 38M-row base — 0.77 test AUC, 8x lift in the top 1% — scoring 2.9M households for a live campaign. The part I'd rather be judged on is what happened next: I found a validation-set defect in my own pipeline (early stopping on the test split), quantified its effect across every figure I had already reported, restated them, and added a pure-noise regression test that pins the model to chance when fed random features — so that class of leak cannot come back quietly. NLP is hands-on rather than API-deep: my degree thesis fine-tuned BERT, RoBERTa and XLNet to state of the art on the FNC-1 stance-detection benchmark, published at LREC 2020.

Building on my own time: github.com/vslovik/fenix — it ranks an incoming stream against a free-text description of what you're looking for, and answers questions over the same corpus with citations back to source chunks. Ollama embeddings, sqlite-vec, no API keys. The part worth looking at is the evaluation: the ranking anchor is scored against a labelled probe set with a deliberate control group of things I don't want, and live results are rated blind — scores hidden, order shuffled — so the human judgement stays independent of the ranking it is judging. Doing that produced a measured finding I did not expect: an embedding has no notion of negation, so naming a technology in order to reject it moves the anchor toward it. Numbers and method in lessons/embedding-anchors.md.

Also a tool-calling agent that turns unstructured regulatory text into a deterministic calculation pipeline — the model does the extraction, a deterministic engine does the arithmetic.

Looking for agentic AI/LLM engineering, LLM evaluation and observability, AI integration, or software architecture. Founding-engineer shape suits me — early employee, not co-founder, but early enough to be in the room where the work gets defined. Direct with the company that owns the product: not consultancy, not agency placement, not a body on someone else's engagement.