Six lines. Everything below is the reasoning.
People use these three titles loosely and often interchangeably. They are genuinely different jobs.
Answers questions about what already happened, and turns the answer into a decision. Given a table of orders, an analyst works out which channel is actually producing revenue, which segment is churning, and what should change next month.
The core skill is not maths. It is asking the right question, knowing whether the data can honestly answer it, and explaining the answer to someone who will not read the spreadsheet. Tools: SQL, spreadsheets, a dashboard tool, some Python.
Builds something that predicts or explains, rather than describing what already happened. Forecasting tomorrow's electricity demand is data science. So is estimating which of your leads is most likely to convert, and by how much.
The core skill is statistics: knowing whether an effect is real or noise, building a model, and honestly measuring how wrong it is against a dumb baseline. Tools: Python, statistics, machine learning.
Builds and maintains the plumbing that makes the data trustworthy in the first place. Nobody can analyse or model their way out of data that was recorded wrong.
The core skill is software and systems: pipelines, schemas, scheduling, making sure the numbers arrive complete and on time. Tools: SQL, Python, orchestration tools, cloud databases.
| Analyst | Scientist | Engineer | |
|---|---|---|---|
| Question | What happened, and so what | What will happen | Can we trust the numbers |
| Output | A recommendation | A model and its error | A pipeline that keeps running |
| Hardest part | Asking the right question | Statistics and honest evaluation | Reliability and schema design |
| Maths load | Light | Heavy | Light |
| Software load | Light | Medium | Heavy |
| Your exposure | Highest, this is Salt work | The IESO thread | You have done more than you think |
An honest read, matched to what you can claim in a live interview rather than what looks good on a page.
In order. Tick a step to mark it done; the ticks are stored on this device only.
You started this and you are keeping it, deliberately, to confirm there are no holes in the fundamentals and to pick up the canonical names for things you already do by instinct. That is a sound reason. Two caveats so it does the job you want it to.
First, it closes AI usage gaps, not data gaps. It is not a data course, so it cannot answer the question "do I have holes in data." Courses 6 and 7, AI for Data Analysis and AI for App Building, are the only two touching this page's subject.
Second, timebox it. Coursera rates it beginner with no prior AI or coding experience required, and lists roughly an hour per course. If it is taking materially longer than that, the reason is worth knowing.
This is the gate. Step 3 is rated Advanced and explicitly assumes SQL, and your own honesty rule says SQL is currently AI-assisted for you. It is also the fastest gap on this page to close, because you are not learning a concept, you are learning syntax for things you already understand.
Do it against your own databases rather than a course sandbox: the location timeline SQLite, the Signal Mapper D1, the Labs `labs` database. Write the query before you ask for help, then compare. The test for "done" is that you can write a join with a group-by and a window function from memory and explain what each does.
This is the one that closes the gap that actually cost you something. Seven courses, roughly 155 to 180 hours, Python and Jupyter throughout, which is the language you write unassisted.
| Course | Hrs | Verdict |
|---|---|---|
| Foundations of Data Science | 20 | Move fast, mostly framing |
| Go Beyond the Numbers | 28 | Worth it, this is the analyst skill |
| The Power of Statistics | 31 | The gap. Do not rush this one |
| Regression Analysis | 28 | Named in the IESO posting |
| Nuts and Bolts of Machine Learning | 34 | Trees and forests, model evaluation |
| Capstone | 6 | Replace the sample data with your own |
| Accelerate Your Job Search with AI | 6 | Skip |
Do not take the beginner Google Data Analytics certificate first just to earn the prerequisite. It is another nine courses and roughly 180 hours, largely spreadsheets and dashboard work you would find slow. Step 2 buys you the same entry for a fraction of the time.
Regression in step 3 gets you most of the way, but forecasting over time has its own toolkit: seasonality, autocorrelation, ARIMA. Demand forecasting is exactly this. Two routes, and the choice is a real tradeoff.
Practical Time Series Analysis from SUNY is the well regarded one, intermediate, roughly 26 to 30 hours, and it covers stationarity, ARIMA, SARIMA and seasonality properly. The catch is that it teaches in R, not Python. The concepts transfer completely, the syntax does not.
If you would rather stay in Python there are shorter applied forecasting courses in the catalogue. They are thinner on theory. Given your background, my read is that the R course teaches you more and the language cost is smaller than it sounds, but this is a genuine judgement call and either is defensible.
Sixteen courses. It is on this page for completeness, not as a recommendation. Two thirds of it is tooling you would only meet inside a company that already runs that stack, and you have already done real engineering work without it. Revisit only if pipeline building becomes the job you are chasing rather than a means to an end.
Coursework without a project produces a certificate. Coursework with a project produces something you can point at. Both of these use data you already own.
The problem is documented and specific. Deals are created at enrolment instead of on reply, so more than a thousand deals sit in stages nothing has ever exited, eleven ever reached Qualified, and none closed. Stage conversion is therefore unmeasurable by construction.
Reply-to-won conversion is the variable that swings the proposed commission by roughly a factor of three. Until it is measurable, the commission cannot be sized, and an unsized commission does not get underpaid, it gets uncalculated.
Why it belongs on a learning page: it is a schema and instrumentation problem, which is data engineering, followed by a conversion analysis, which is analytics. You would be doing the two roles in sequence on data you already have access to, with a financial payoff attached.
Already specced in your own notes. IESO publishes hourly Ontario demand openly, so there is no confidentiality problem and no permission to ask for. The build is: pull the data, explore the daily and seasonal shape, build a naive baseline, build a regression on calendar and temperature features, then score both honestly on a held-out period.
The reason this is good practice rather than a toy is the baseline. Being able to say "my regression beat seasonal-naive by this much on MAPE, and here is the kind of day where it falls apart" is a real answer. It is also the same methodology as gating a learned model against gradient boosted trees, which you have already practised.
Sequencing: the statistics and regression courses in step 3 make this project much better. Do a rough version first anyway, then redo it properly afterwards. The gap between the two versions is the clearest evidence you will have that the coursework worked.
Coursera Plus, monthly, in Canadian dollars, renewing on the 8th. It is not a subscription to one program. It is access to essentially the whole catalogue, which means every certificate on this page is already paid for. There is nothing to buy. The only currency this page spends is your evenings.
There is a switch-to-annual offer on your purchases page that lowers the effective monthly cost meaningfully. It is paid upfront for the year, so it only makes sense once step 3 is genuinely underway and you know you will still be here in six months. Revisit it then, not now.