Most workflow posts list steps. This one follows one small table through them: my Kaggle dataset of reported serious crimes in Guyana, 1990 to 2022.
Know where the numbers come from
The figures come from the Criminal Investigation Department of the Guyana Police Force, as collected by the Bureau of Statistics. They count crimes that were reported, which is a different thing from crimes that happened. Write that down before you draw a chart.
The file has 33 rows, one per year. There is a column for each serious crime: murder, manslaughter, wounding with intent, burglary and break-in, larceny, arson, rape, and other. "Other" covers offences such as perjury and escape from lawful custody. The last column is the total.
Load it and look
Small files still need cleaning. This one has four problems.
- Two headers have line breaks inside them, so "Wounding with Intent" loads as a name split over two lines.
- Most counts carry a thousands comma, like "3,222". A few recent ones do not, like "1316". The column loads as text until you fix that.
- A dash means the figure was not available. There are two: arson in 2010 and wounding with intent in 2011.
- Wounding with intent is 0 in 2022, after 158 the year before. That could be a real zero or a missing value written as one. The file does not say which.
Check the arithmetic
The table carries its own total, so test it. In 32 of the 33 years, the categories add up to the total. In 1994 they add up to 5,111 and the total says 5,188. Either a category is missing or a number was copied wrong, and the file cannot tell you which. Flag it and go back to the source.
Then look for trends
Burglary and break-ins fell from 3,222 in 1990 to 666 in 2022. Larceny jumps from 2,898 in 2013 to 12,639 in 2014, then drops to 5,171 in 2015, and that one year pushes the total to 14,855. Before anyone explains that spike, someone should check that 2014 counted larceny the same way as the years around it.
Model last
The dataset asks one question in its subtitle: can you find trends and patterns in crimes over the years? With 33 rows, a line chart per category is a good first answer. Skipping the checks above to start modelling is tempting when a deadline presses. It is also how projects produce confident wrong answers.
Ship it with its notes
A model that works in a notebook but never reaches the person who needs it has not finished the workflow. The dataset is public on Kaggle under the CDLA-Sharing 1.0 licence, with its source in the description, so anyone can check it against the Bureau's table.