Skip to content
SK

AI Lab

Exploring the future of intelligent systems

The areas I work in, experiment with, and am still learning — labeled honestly. Some of this is shipped in production. Some of it is a weekend project. Some of it I am only reading about, and it says so.

Applied in production work

Used in shipped professional or client work.

Hands-on

Built with it in self-directed projects.

Exploring

Actively learning. Not yet shipped.

Applied in production work

Shipped in production work

Used in shipped professional or client work.

Agent Optimization

applied

Making multi-agent systems reliable enough to trust: how work gets delegated, what each agent remembers, and how you tell whether a change actually helped.

Tool callingAgent planningOrchestrator / sub-agent designShort-term, long-term & episodic memoryContext engineeringModel routingAgent evaluationReliability

RAG & Knowledge Systems

applied

Retrieval past the naive baseline - where chunk similarity stops being enough and the structure of the data has to be modeled directly.

Metadata-aware chunkingHybrid searchBM25Graph RAGParent-document retrievalHyDERerankingRetrieval evaluation

AI Security & Guardrails

applied

What happens when someone tries to talk your agent out of its instructions - and what stands between a model and the data it should never repeat.

Prompt injectionJailbreak resistancePII / PHI screeningInput & output guardrailsSecure tool callingData privacyAI safety

LLM Evaluation

applied

Turning “the output looks better” into a number you can regress against. Judge harnesses, labeled datasets, and shadow evaluation off the serving path.

LLM-as-judgeLabeled query/response datasetsFaithfulness & relevancy scoringShadow evaluationBayesian threshold optimizationAccuracy SLAs

Graph & Query Systems

applied

When relationships are the answer, not the metadata. Modeling domains as graphs and letting people ask in plain language.

Neo4jAWS NeptuneCypherGremlinNatural-language-to-CypherNatural-language-to-SQLSchema introspectionTopic modeling

Core ML, NLP & Vision

applied

The foundation underneath the LLM work - the classical modeling that still decides whether a pipeline is any good.

Machine LearningDeep LearningNLPComputer VisionPyTorchTransformersScikit-learnTensorFlow

Hands-on

Built in self-directed projects

Built with it in self-directed projects.

LLM Training & Fine-Tuning

hands-on

Fine-tuning smaller models for the jobs that do not need a frontier model - classification, routing, and structured extraction.

TransformersFine-tuningDeBERTa-v3Synthetic dataset generationSFTPEFT / QLoRAQuantizationModel evaluation

AI Infrastructure & Cost

hands-on

The part nobody demos: what a system costs per query, where the latency goes, and how to route around both.

Inference gatewaysResponse cachingCost-tiered routingLatency optimizationDockerMLflowOptunaAWS (EC2, S3, SageMaker, Bedrock)

Exploring

Actively learning

Actively learning. Not yet shipped.

Multimodal & Generative AI

exploring

Systems that take in more than text. Currently a reading and prototyping area rather than shipped work.

Multimodal modelsSpeech interfacesVision-language modelsGenerative AI patterns

Working on something in one of these areas?

I take on a small number of freelance engagements, and I am always up for comparing notes on agent architecture and retrieval design.