"Things go wrong. The odds catch up. Probability is like gravity: you cannot negotiate with gravity." James 'Sonny' Crockett, Miami Vice The Movie.
Monday, June 30, 2025
Sunday, June 29, 2025
What is a Shock Decomposition?
The Nixon Shock of 1971 was an attempt to stop inflation by ending convertibility of the US Dollar into Gold. Along with other forms of economic Shock Therapy, is (from the standpoint of Systems Theory) an attempt to pop Economic Bubbles and return the system to a sustainable attractor path. But economies are subject to all sots of shocks and the general question for analysis is how do feedback effects operate in the presence of shocks. The way to study the effects of shocks on a system is called Shock Decomposition And, if you have a computer model, shock decompositions can be studied through computer simulation.
What a shock decomposition will show you depends on (1) what parts of the system are being shocked and (3) the internal dynamics (if any) of the computer model. To focus this discussion, I am going to concentrated on systems and state space models since any system can be put in state space form (see the discussion here).
Take a very simple state space (SS) model:
Thursday, April 17, 2025
World-System (0-2000) The Maddison Database
The most complete data set on the World-Economy is the one developed by Angus Maddison and published by the OECD (here). I have entered the data into seven spreadsheets for Population, GDP, Urbanization, Real Exports, Exports, Total Hours worked and Employment. The tables are reproduced below.
To make continuous series available, I have nonlinearly interpolated missing data using the Spline Smoothing algorithm in the R programming language. Where initial data is missing, I have used the E-M Algorithm to estimate missing values. In some cases for some countries and regions, no data is available.
Boiler Plate
State Space Model Estimation
The Measurement Matrix for the state space models was constructed using Principal Components Analysis with standardized data from the World Development Indicators. The statistical analysis was conducted in an extension of the dse package. The package is currently supported by an online portal (here) and can be downloaded, with the R-programming language, for any personal computer here. Code for the state space Dynamic Component models (DCMs) is available on my Google drive (here) and referenced in each post.
Atlanta Fed Economy Now
My approach to forecasting is similar to the EconomyNow model used by the Atlanta Federal Reserve. Since the new Republican Administration is signaling that they would like to eliminate the Federal Reserve, the app might well not be available in the future.
While the app is still available, there have been some interesting developments. In earlier forecasts, the Atlanta Fed was showing GDP growth predictions outside the Blue Chip Consensus. Right now, after unorthodox economic policies from the Trump II Administration, the EconomyNow model is predicting a drastic drop in GDP (the Financial Forecast Center is only predicting a slight drop here).
Climate Change
Another comparison for what I have presented above are the IPCC Emission Scenarios. These scenarios are for the World System. Needless to say, (1) the new Right-Wing Republican administration plans on withdrawing the US from all attempts to study or ameliorate Climate Change and (2) the IPCC does not produce any RW modes for the World System (but seem my forecasts here).
World System
The longest running set of data we have for the World-System is the Maddison Project based on the work of Angus Maddison (more information is available here). Data on production (Q) and population (N) for most countries and regions runs from years 0-2000. More data becomes available as we near the year 2000.
Available data were entered in a spreadsheet (see Population above, double click to enlarge). Missing data were interpolated with nonlinear spline smoothing using the R programming Language.
In cases where initial values were not available (see GDP above), the E-M Algorithm was used to estimate initial conditions.
From the graph of GDP above (W_Q) for the World System, it can be seen that economic growth from the year 0-1500 was basically flat. The period of British Capitalism (after 1500) had a small plateau of growth. Takeoff does not happen until the Nineteenth Century.
From a system's perspective, the only model that can be tested for the entire period is Kenneth Boulding's Malthusian Systems Model [Q,N] = f[Q,N].
When developed as a State Space model (measurement matrix above) there are two components: W1=Growth and W2=(Q-N), the Malthusian Controller. When more data is available, the Malthusian Controller can be generalized to other SocioEconomic theories.
What the Malthusian Controller shows (plotted as Q-N above) is that a long-developing Malthusian Crisis (Q<N) started in the Late Middle Ages and accelerated through the period of British Capitalism (Dark Satanic Mills) and was reversed spectacularly during the Nineteenth Century. Takeoff in response to a deepening Malthusian Crisis would not be an unreasonable way to view Modern Economic Growth.
Error Correcting Controllers (ECC)
In another post (here), I presented Leibenstein's Malthusian Error Correcting Controller (ECC). It can be generalized to the dominant ECCs in most theoretical economic models (above). These controllers can be further generalized. For example, (X-U) and (L-U) can be generalized to (N-U), a more general Urbanization Controller which describes market expansion for economic growth. In countries and periods with limited data, (N-U) might subsume all these processes. ECCs describe important feedback processes in SocioTechnical System that are typically not recognized as such in academic literature.
Kaya Identity
The basic theoretical model underlying all the World-System models I create is the Kaya Identity. There are a number of advantages to starting theoretical development with the Kaya Identity: (1) An "identity" is true by definition Adding other variables to the model ensure that theory construction is on a solid footing. (2) The Kaya Identity is used as the basis for the IPCC Emission Scenarios.
World Development Indicators (WDI)
After WWII, extensive data sets on all countries in the World-System became available from the World Bank (here). The indicators above where chose to construct the state space for each WDI-based model. Addition indicators can be added for specific forecasts and analyses.
Wednesday, February 5, 2025
Dynamic Component Models (DCMs)
The Dynamic Components Model (DCM) is a form of State Space Model where the approximate state variables are first computed using Principal Components Analysis (PCA). The difference between a standard State Space Model and the DCM are (1) how the Measurement Matrix (see graphic above) is computed (with PCA) and (2) how the state variables are analyzed directly. The advantage of the DCM is that it separates the growth and control state variables so that growth and control can be analyzed independently (the PCs are statistically independent). In the Cannonical DCM, there are only three state variables (one for growth and one for control) that explain at least 80% of the variation in the Output Variables. The control components are typical made up of Error Correcting Controllers (ECCs). In SocioEconomic systems the ECCs are typically associated with a theoretical tradition. For example, the Malthusian Controller is presented here and generalized here. ECCs are key elements of Cybernetics and the dominant growth component is a key element of Systems Theory and Economic Growth Theory.
The DCM is implemented in the public domain R Programming language as an extension to the dse (Dynamic System Estimation) package. The dse package can be downloaded on all Computer platforms and can be run on line with a web browser (here). The DCM extensions with documentation are available here.
In the dse package, a state space model can be created using the SS command in R:
Properties of State Space Models
Stability
Mechanization
Controllability
Reachability
Stability
McMillan Degree
Monday, December 30, 2024
Forecasting and Scenario Construction
Of course, we cannot know The Future. Think about trying to predict the Future in 1900: World War I, The Great Depression, World War II, the Cold War, etc. Even after the fact, we have trouble explaining what happened in the InterWar Period. Yet, we persist. We try to predict the effects of Climate Change. We try to predict the path of Hurricanes. We try to predict next Quarter's Economic performance.
Our vision for all this effort is (1) Science Fiction Psycho-Historian Harry Seldon created a hand-held device called the Prime Gradient that predicts the collapse of the Galactic Empire in the Foundation Trilogy and (2) Limits to Growth Engineer J. Wright Forrester created a computer program, World1, that predicted the collapse of the World System in 2050 as a result of resource shortage. These aren't the work of cranks. Issac Asimov was a scientist. J. Wright Forrester taught at MIT.
It is a mistake to think anyone knows the future. It is unknowable. What I think we can do is Explore the Future with state space models, systems analysis, multi-model inference and scenario construction. Some scenarios constructed in this manner are too awful to contemplate and must be avoided at all costs. An entire group of middle-range scenarios will seem likely but the best model cannot be chosen in advance. Systems are too complex and there is significant random error. An example of my approach is Five Futures for Russia in which I use five models of the Russian SocioEconomic system to construct statistical scenarios for the future (one is a statistical surprise). You can actually run the Business as Usual Model (BAU) on line here.
Computer simulation of statistically estimated systems models is essential. We have to get beyond the stage of arm-chair speculation which still seems to be the privileged mode of academic discourse based on The Classics. When faced with having to make predictions about the future of Climate Change, the science-based IPCC made the right choice: simulation and scenario construction. The Social Sciences have supplied few useful models for the IPCC project. Neoclassical Economics has provided the DICE Integrated Assessment Model but it is based on flawed neoclassical assumptions and is not statistically estimated or tested.
Here is some more detail on my approach. The methodology is all readily available and it remains for the Social Sciences to make a serious effort to apply it.
Atlanta Fed Economy Now
Hurricane Forecasting
Climate Change
Thursday, December 10, 2020
Confounded COVID Vaccine Trial
Thursday, September 12, 2013
A Flowchart For Quibbling With Research Results
Friday, December 28, 2012
A Better Black Friday Retail Sales Model
The analysis of aggregate retail sales is essentially a time series problem in that we are analyzing sales over time for, ideally, multiple years. HLMs can be written to handle time series problems, a topic that I will return to in a later post. The analysis of aggregate US retail sales, however, can be analyzed with a straight-forward state space time series model. If you are interested in the pure time series approach, I have done that in another post (here).
The insight from the time series analysis is that US Retail Sales are being driven by the World economy which makes sense in a world of globalized retail trade. The idea that Black Friday Sales might be a good predictor of aggregate retail sales is, based on purely theoretical considerations, not a very appealing hypothesis.
Friday, December 7, 2012
How to Write Hierarchical Model Equations
First, I will review the examples I have already used and in future posts introduce a few more. When I develop equations, there is an interaction between how I write the equations and how I know I will have to write simulation code to generate data for the model. Until I actually show you how to write that R code in a future post, I'm going to use pseudo-code, that is, an English language (rather than machine readable) description of the algorithm necessary to generate the data. If you have not written pseudo-code before, Wikipedia provides a nice description with a number of examples of "mathematical style pseudo-code" (here).
For the Dyestuff example (here and here) we were "provided" six "samples" representing different "batches of works manufacturer." The subtle point here is that we are not told that the batches were randomly sampled from the universe of all batches we might receive at the laboratory (probably an unrealistic and impossible sampling plan). So, we have to deal with each batch as a unit without population characteristics. So, I can start with the following pseudo-code:
For (Every Batch in the Sample)
For (Every Observation in the Batch)
Generate a Yield coefficient from a batch distribution.
Generate a Yield from a sample distribution.
End (Batch)
End (Sample)
If we had random sampling, I would have been able to simply generate a Yield from a sample Yield Coefficient and a sample Yield distribution (the normal regression model). What seems difficult for students is that many introductory regression texts are a little unclear about how the data were generated. On careful examination, examples turn out to not have been randomly sampled from some large population. Hierarchical models, and the attempt to simulate the underlying data generation process, sensitize us to the need for a more complicated representation. So, instead of the single regression equation we get a system of equations:
The second model I introduced was the Black Friday Sales model (here). I assumed that we have yearly weekly sales data and Black Friday week sales generated at the store level. I also assume that we have retail stores of different sizes. In the real world, not only do stores of different sizes have different yearly sales totals but they probably have somewhat different success with Black Friday Sales events (I always seem to see crazed shoppers crushing into a Walmart store at the stroke of midnight, for example here, rather than pushing their way into a small boutique store). For the time being, I'll assume that all stores have the same basic response to Black Friday Sales just different levels of sales. In pseudo-code:
For (Each Replication)
For (Each Size Store)
For (Each Store)
Generate a random intercept term for store
Generate Black Friday Sales from some distribution
Generate Yearly Sales using sample distribution
End (Store)
End (Store Size)
End (Replication)
and in equations
where the terms have meanings that are similar to the Dyestuff example.
I showed the actual R code for generating a Black Friday Sales data set in the last post (here). The two important equations to notice within the outer loops are
b0 <- b[1] - store + rnorm(1,sd=s[1])
yr.sales <- b0 + b[2]*bf.sales + rnorm(1,sd=s[2])
The first gets the intercept term using a normal random number generator and the second forecasts the actual yearly sales using a second normal random number generator. The parameters to the function rnorm(1,sd=s[1],mean=0) tell the random number generator to generate one number with mean zero (the default) and standard deviation given s[1]. For more information on the random number generators in R type
> help(rnorm)
after the R prompt (>). In a later post I will describe how to generate random numbers from any arbitrary or empirically generated distribution using the hlmmc package. For now, standard random numbers will be just fine.
In the next post I'll describe in more detail how to generate the Dyestuff data.
Thursday, December 6, 2012
Generating Hierarchical Black Friday Data
Since this posting is meant to show you how to generate your own Black Friday data set, load the hlmmc libraries and procedures as usual in the R console window (using the hlmmc package is described here and the R programming language, a free multi-platform public-domain program, is available here):
To make things simple to start (I'll explain how to make it more complicated in a future post), assume that the red store lines are all parallel. In terms of a regression equation, the only difference between these lines then is in the constant term.
The first equation above models that term. (Store) is an indicator variable which is zero for the largest store (imagine it's Target). Lambda_00 plus some normally distributed error term, mu_0j, determines the intercept value for that store. To create data that looked like Paul Dales', I chose a value of 1.3 for lambda_00. The smallest store (imagine a small, hip, boutique shop) was coded 3 since it was the fourth type of store by size and would have an intercept that was 3 less than Target plus some error term. Stores of other sizes would fall in between. The regression lines would all be the same with a value chosen at Beta_1j = 0.6 in the second equation.
I chose the Black Friday Sales value (the X in the second equation) by generating a uniform random number between -2 and 0 (the approximate values for the biggest store, say Target) and then shifted the value upward based on the store number (the details of this are presented in the code below). The modeling assumes that, as part of our sampling plan, we deliberately chose stores of different sizes (again, I have no idea how the original stores were sampled by Mr. Dales but the simulation exercise is designed to make me think about the issue).
The third equation above is the naive regression equation used in Paul Dales' analysis. By substituting in our more realistic data generating model in the fourth equation we see that we actually have two error terms, mu_0j and epsilon_ij. This result should sensitize us to the possibility that we might not find significant results in a naive regression model if the data were really generated hierarchically with error both at the store- and aggregate-levels .
Hopefully, the data generating model is now clear and we can start creating some simulated data. In the R console window you can type the following (after the prompt >):















