Company · Consulting

Data & research for hire

The same data engineering and quantitative research discipline behind DoubleTrends™, available for custom projects. We help teams turn messy data, market questions, and forecasting problems into reproducible analysis and usable systems. Scope and pricing depend on complexity — get in touch to discuss your needs.

Process

From research question to working deliverable

01 · Scope

Define the business question, available data, constraints, timeline, and what a useful answer should look like.

02 · Build

Clean the data, run the analysis or model, document assumptions, and review intermediate findings before final delivery.

03 · Deliver

Provide code, datasets, charts, methodology notes, and a short explanation of how to maintain or extend the work.

Machine Learning illustration
Machine Learning

Machine Learning Model Development

Model development for structured tabular data: classification, regression, time-series forecasting, and anomaly detection. Common problems include churn prediction, demand and revenue forecasting, risk scoring, fraud detection, and customer segmentation. Work covers data review, feature engineering, model selection, training, and rigorous evaluation on a held-out test set — with intermediate findings reviewed before final delivery, not just at the end.

Deliverables include trained model artifacts, reproducible Python code, and a model card documenting feature importance, performance by segment, known failure modes, and a plain-language statement of what the model can and cannot be trusted to do. Most projects run three to eight weeks depending on data readiness and problem complexity. You should come with a structured dataset and a defined outcome variable or a clear business question; scoping starts from there.

Forecasting illustration
Statistics & Forecasting

Statistical Analysis and Forecasting

Exploratory data analysis, hypothesis testing, regression modelling, and time-series forecasting using R or Python. Specific techniques include t-tests, ANOVA, chi-square tests, linear and logistic regression, difference-in-differences, regression discontinuity, and time-series decomposition (STL/SEATS). Suitable for causal inference questions, A/B experiment evaluation, metric attribution, and building forward-looking models from historical data.

R is preferred when the primary output is a statistical report or a presentation-ready set of charts; Python when the analysis feeds downstream code or a data pipeline. Results are delivered with explicit methodology documentation — assumptions stated, confidence intervals given in plain language, and a clear statement of what the analysis can and cannot conclude. The goal is output that supports a real decision, not a notebook that sits unread.

Data mining illustration
Data Collection

Data Collection and Scraping

Automated data collection from public web sources, APIs, and structured documents. Technical scope includes JavaScript-rendered pages (via Playwright or Puppeteer), REST and GraphQL API integration, and document parsing from PDFs and structured reports. Output is a clean, validated dataset — schema-documented, type-enforced, and deduplicated — delivered in your preferred format: CSV, JSON, or direct database write.

Typical use cases include market data aggregation, competitor pricing intelligence, public financial data pipelines, and building training datasets for downstream modelling. For ongoing collection, the pipeline includes scheduled runs, error logging, and alerting; for one-off work, you receive a documented snapshot with enough notes to reproduce the collection if needed. Work is limited to public sources where automated access is consistent with the source’s terms.