Search This Blog by Google

Search This Blog

Welcome to Dijemeric Visualizations

Where photography and mathematics intersect with some photography, some math, some math of photography, and an occasional tutorial.

Total Pageviews

Showing posts with label statistical. Show all posts
Showing posts with label statistical. Show all posts

Saturday, January 17, 2015

Random Events, Trends, and Hot Temperatures


Random Events, Trends, and Hot Temperatures
(c) Ken Osborn Jan 17, 2015



Today's papers (Jan 17, 2015) say 2014 was the hottest on record since 1880. An engineering friend says the trend may be just a random thing. Could that be? Though the data do show a reasonable closeness to an independently produced set of random values with the same average and variability (standard deviation), I decided to test it by comparing the first half of 1880-2014 to the second half.

If random, dividing the data into two halves and ordering the values from lowest value to highest within each half would produce two superimposed sets that would be indistinguishable. If they are part of a trend of increasing temperatures, the sets would not superimpose and the second set would be shifted to the right on the chart. So which is it? Check out the graphs.
 [Data source: http://data.giss.nasa.gov/]


First, let me state that I do understand that random events can generate what appears to be a trend, as in this example:




Chart 1: Example of a 'trend' from a random generation of values 


The Random Walk chart displays what appears to be a strong trend with a correlation coefficient of 0.96, sufficient to earn an award for superior performance by an inebriated soul staggering along a mostly straight path.  I should note that it took over 50 runs of the program to get this result and doubt I could reproduce this particular set again.  


But if we can get a randomly generated collection of results to look like a trend, how do we know that our trend was not generated randomly?  

One way, of course, would be to rerun the events and see if we get the same trend.  However, since I'm looking at data collected from 1880 to 2014 that is not a realistic approach.  Another way would be to take the data, split it into two halves, and compare the two halves.  If the data are a random distribution around a central value, the first half of the data should match the second half of the data.  But first, let's try that with some randomly generated data.



Chart 2: Plot of 100 random values generated by a Monte Carlo simulation 


In chart 2, 100 random values ordered from low to high are plotted against their rank (probability). Half of the values exceed and half are below the average of zero and are symmetrically distributed around the center.  In other words, these data exhibit a nice Normal type distribution.  Note that the average and standard deviation for these data are the same as the temperature data to follow.  




So now if those data are split into two halves and each half is ordered independently from lowest value to highest value and plotted against its rank, what do we get?  



Chart 3:  Comparison of two halves of a randomly generated set of values after ranking each half


The values in blue represent the first half of 100 random values and the ones in red the second half.  Each half was then ranked from lowest to highest value and plotted against the ranking (1 to 50).  The superimposition is not exact, but these are randomly generated results so one should not expect an exact agreement between individual values.  But maybe a real trend would also show something like this.  Let's try it.  



Chart 4: A trend (Y= 10X+5) of 100 values partitioned into two ordered halves 


The trended data set of 100 values was generated from the formula Y = 10X+ 5.  The set was divided into two halves and each half was ordered from lowest value to highest value then plotted against its rank from 1 to 50.  Unlike in chart 3 with conformable data sets, these two sets show no overlap at all.  So how will this work with real data?  


 

Chart 5: Comparing the ranked temperature anomalies from 1880 to 2014 with a random data set  (Source: http://data.giss.nasa.gov/)

The values in green in chart 5 represent the temperature anomalies from 1880 to 2014 ranked from lowest to highest value.  Each value represents the deviation from the average for the 20th century.  The values in blue were randomly generated using the mean and standard deviation from the temperature anomaly set.  They do look as close as two separate runs of a random number generator.  But remember, the real test is to see if the first half of the data (1880 to 1946) matches the second half (1947 to 2014).  Any guesses?  


Chart 6: Comparison of two halves of the 1880-2014 temperature data anomalies


In chart 6, the values in red are for the years 1880 to 1946 and the values in green for 1947 to 2014.  Each set is ordered from its lowest to highest value and plotted against corresponding year.  They do not match and are clearly two separate distributions.  I leave the conclusion to you as to whether these data have been generated by random events.    


Tuesday, January 31, 2012

Predator vs Prey - A Mathematical Model - Part 2 of a 3 Part Series

In the first installment of my predator-prey model I presented graphs of population growth under three scenarios that included deer and forage but no predators.  Scenario 1 was for a stable environment (i.e, constant carrying capacity): the deer herd rapidly grew from a starting population until it reached the carrying capacity then leveled out.  Scenario 2 was a moderately variable environment and the deer initially grew but the population fluctuated above and below the carrying capacity.  Scenario 3 was for a highly variable environment.  In this last case, the deer herd grew, fluctuated in number, then crashed and died out.  For a more complete review see http://misterkenblog.blogspot.com/2011/12/predator-vs-prey-will-wolves-dominate.html.

What happens when predators are introduced?  Will the deer herd die out sooner or will it stabilize because wolves keep the deer herd in check with the environment?  Can predictions even be made?  Let's see.

Scenario 1: Start with a deer herd well below the carrying capacity, a small number of wolves, and a constant environment.



The starting conditions are an initial deer population (N) of 5000, a growth rate (R) of 0.5(50% increase in deer herd per generation), carrying capacity (K) of 20000, no variance in the carrying capacity (KV=0), 5 wolves, a reproductive capacity of 0.1 (10%), and a predator efficiency (E) of 0.53 (53% - VERY good hunters), and a requirement (S) of 24 deer/wolf/year to sustain wolf pack growth.

Both the deer herd and wolf pack show rapid initial growth, followed by oscillations of population size, and finally a steady state based on the carrying capacity for the deer herd.  The final steady state is well above the starting conditions for both deer and wolves and the final state for the deer is about half of the carrying capacity.
Predation in Stable Environmet



Scenario 2: Add global warming or some other factor to make the environment variable



The variance for the carrying capacity has been increased from 0 to 1000, or a variance factor of 20% (100x1000/5000).

The oscillations in both deer and wolves has increased and no steady state is achieved although after 500 generations it does not appear that neither deer nor wolves are in danger of extinction.

Add Moderate Amount of Environmental Variability


Scenario 3: Scenario 2 with more predation by increasing the wolf pack from an initial 5 to 50



Increasing the initial number of wolves from 5 to 50 does not change the overall dynamics.  It would appear that starting with an initial wolf pack of fewer than what is sustainable has little long term effect.  I leave it for the reader to try other starting wolf pack sizes once I have posted the interactive spreadsheet.

Increase Predation Pressure


Scenario 4: Keep the starting wolf population at 50 and increase the environmental variability


Increasing the variance in carrying capacity to 100% dramatically shifts the oscillations in both the deer herd size and wolf pack numbers. While neither population crashes, they come perilously close.  I leave it to the reader to try more simulations to see if the wolves, or deer, or both go to extinction under these conditions.

Increase Environmental Variability



Scenario 5: Same as scenario 4 with initial carrying capacity cut in half


In this last scenario, the deer herd crashes and the wolves, lacking a food supply, follow.

Reduce Carrying Capacity

Whether the wolves control the deer population or the deer control the wolf population is still an open question, but clearly the environment controls both.  When the deer herd exceeds the carrying capacity of the environment extremes from one year to the next will ultimately result in a population crash.  In the absence of predation, by this model, the deer herd will subsist only if the environment is very stable.  If the environment is not stable (the normal course of events) predation pressure can help stabilize the deer herd by reducing the herd size and the odds that the deer herd will not exceed the carrying capacity are improved.  Of course if the deer herd crashes so will the predators unless they have a reserve food source.  That in fact is the case, but then it becomes a matter of energetics and whether switching to an alternative food supply for the predator is analogous to a drop in carrying capacity for the prey.


 Next month I will post a link to the statistical model so that the reader can try some scenarios and draw their own conclusions.