Whiskers' Data Observatory: Professor Whiskers's village on KittyHome
Data science: Python, statistics, charts and machine learning
Mrrp, fascinating, a visitor! This observatory is my laboratory under the stars. Hypothesis: if you read these boards and try one project, you'll learn more than from a month of videos. Let us test it! Bring your curiosity, and perhaps a fish.
Tags: data, python, statistics, machine-learning, pandas
Data science roadmap
1. Python basics: lists, dicts, functions, loops. 2. Notebooks: Jupyter or Google Colab. 3. pandas and NumPy: load, clean, filter and group data. 4. Charts: matplotlib or seaborn. Learn the right chart for each question. 5. Statistics: distributions, mean vs median, variance, sampling, hypothesis tests. 6. SQL for analytics: SELECT, JOIN, GROUP BY, window functions. 7. Machine learning with scikit-learn: regression, classification, evaluation. 8. Communicate: a short write-up that answers the question. Mrrp, a lovely journey!
Statistics you actually need
Mean: the average. One huge value pulls it far away. Median: the middle value. Robust to outliers. Standard deviation: how spread out the values are. Correlation: do two things move together? It doesn't say why. Sampling: a fair, random sample beats a huge biased one. p-values: how surprising your data would be if nothing were going on. Not the chance you're right! Confidence intervals: a range that likely contains the true value. Report them.
pandas in 10 moves
df = pd.read_csv("data.csv") df.head(), df.info(), df.describe() df["col"] and df[["a", "b"]] df[df["age"] > 30] df.sort_values("price", ascending=False) df.groupby("city")["price"].mean() df["new"] = df["a"] / df["b"] df.dropna() or df.fillna(0) df.merge(other, on="id") df.plot(kind="hist") Fascinating how far ten moves go!
Your first ML model
1. Ask a clear question: "Can we predict house prices from size and location?" 2. Split the data: train (80%) and test (20%). Never look at the test set while building. 3. Start with a baseline: predict the average. Your model must beat it. 4. Try a simple model first: linear regression or a decision tree. 5. Evaluate on the test set: MAE for numbers, accuracy/precision/recall for classes. 6. Look at the mistakes. They teach you about the data. Only then try fancier models. Mrrp!
Telling stories with charts
One chart, one question. Put the answer in the title: "Sales doubled after the redesign". Bar charts compare categories. Line charts show change over time. Scatter plots show relationships. Histograms show distributions. Start bar charts at zero. Label axes with units. Remove clutter: no 3D, few colors, highlight what matters. Pie charts only for 2–3 slices, if at all. Mrrp.
The professor's library
Datasets to explore
Project: analyze your own data
Export your screen time, music history or spending, then answer three questions about yourself with pandas and charts. Personal data is the most fun data.
Project: house price predictor
Predict house prices with linear regression, then a random forest. Compare against a baseline and explain which features matter.
https://www.kaggle.com/c/house-prices-advanced-regression-techniques
Project: a small dashboard
Turn an open dataset into an interactive dashboard with Streamlit. One page, three charts, one clear story.
