Data science portfolio · Universitat Politècnica de València

Lianghao Zhan

Reproducible public-data projects in data engineering, urban analytics, forecasting, NLP retrieval and machine learning.

Data Science student seeking an internship in Valencia, hybrid or remote within Spain.

  • 06case studies
  • Public datadocumented sources
  • Reviewablecode, tests and limits
Evidence index · 06 systems

Working principle

Work that can be inspected, not just described.

Each project has a defined question, documented data source, executable pipeline, measured outputs and an explicit boundary around what the output does not prove.

These are portfolio studies over public data. They are not presented as work delivered for companies or as deployed production systems.

Selected work

Six projects built around data decisions.

Filter by domain or search the methods and artifacts used in each case study.

6 case studies visible

Static equity decision dashboard for bicycle parking in Valencia.

Geospatial decision support

Valencia Bike Equity

An exploratory, reproducible system for reviewing territorial equity in bicycle parking across Valencia.

  • Geospatial
  • Urban analytics
  • Decision support

Public data Valencia Open Data

parking points processed
4,316
MCDA weight scenarios
10,000
Static operations dashboard for the Valenbisi snapshot analysis.

Operations research and mobility

Valenbisi Pulse

A reproducible snapshot control center for station risk, local pressure, minimum-cost rebalancing and conservative stress tests.

  • Operations research
  • Mobility
  • Optimization

Public data CityBikes API / versioned local sample

stations in reproducible sample
30
critical stations at base thresholds
21
Benchmark scorecard comparing lexical, semantic and hybrid retrieval strategies.

NLP retrieval and open data

Valencia Open Data Navigator

A hybrid retrieval system for a versioned Valencia CKAN catalog, with benchmarked ranking, API and an interactive discovery interface.

  • NLP
  • Information retrieval
  • Open data

Public data Valencia Open Data CKAN catalog

datasets in versioned snapshot
296
manual relevance queries
20

Portfolio index

A compact evidence ledger.

The same filters above update this table. Values are pulled from committed project runs and link back to the corresponding case study or repository.

Case studyDecision lensMeasured evidenceReviewable artifactsReview
Valencia Bike EquityGeospatial decision supportGeospatial · Urban analytics10,000MCDA weight scenariosOffline snapshot · Spatial diagnostics · MCDA simulationRepository
Valenbisi PulseOperations research and mobilityOperations research · Mobility21critical stations at base thresholdsData validation · Risk scoring · Linear programRepository
Spain Electricity Demand Forecast LabForecasting and data platformForecasting · Data engineering12expanding backtest windowsETL manifests · SQL marts · Temporal backtestsRepository
E-commerce Conversion MLGoverned supervised learningMachine learning · MLOps0.7602deployment-safe ROC-AUCModel pipelines · Ablation study · CalibrationRepository
Valencia Open Data NavigatorNLP retrieval and open dataNLP · Information retrieval20manual relevance queriesCatalog snapshot · BM25 and LSA · RRF and MMRRepository
Valencia Air Quality LakehouseData engineering and SQLData engineering · SQL6.04Mlong-format measurementsSource contract · Dimensional model · SQL quality martsRepository

Evidence ledger

What is visible in the repositories.

Models are not only compared. Their inputs, validation design, uncertainty, operational assumptions and failure modes are surfaced as first-class artifacts.

01

Provenance

Public sources, versioned snapshots and data cards make the origin of each dataset inspectable.

02

Validation

Temporal splits, leakage checks, quality controls and statistical comparisons are attached to the decision context.

03

Reproducibility

Source code, tests, linting, generated reports and CI workflows are available in every project repository.

Technical practice

A portfolio built around decisions.

The tools change with the problem; the working habits stay constant.

01

Data foundations

Python, pandas, SQL, DuckDB, validation checks, manifests and carefully scoped data transformations.

02

Modeling and inference

scikit-learn pipelines, calibrated probabilities, temporal backtesting, bootstrap intervals, spatial diagnostics and optimization.

03

Communication and quality

Plotly and Streamlit exploration, model and decision cards, pytest, Ruff, Docker where useful and GitHub Actions.

Availability

Looking for a Data Science internship.

Open to roles in data analysis, data science, machine learning, business intelligence, automation and junior data engineering in Valencia, hybrid or remote in Spain.