Climate2003 · Original research pages, 2003–2007

THE AUDIT TRAIL

We have attempted to provide a comprehensive audit trail to enable third parties to independently verify as much as possible.  We have included for reference the proxy dataset (1 MB) used by Mann Bradley and Hughes. It can be opened in an Excel sheet (or you can get an Excel version at the review&critique web site. The audit trail for understanding MBH98 errors is divided up into 3 sections:

1.      Errors and defects which can be verified through inspection of the MBH98 dataset

2.      Updates which can be verified through comparison of MBH98 and WDCP data

3.      Errors in proxy principal component calculation, which require re-collation of WDCP data, and comparison of explained variance.

For the final version of the data set after corrections, scroll down to the bottom of the page.  

Note: “ NaN ” or "NA"  is the code for missing values.

<< Return to Main page.

 1. Errors which can be verified through inspection

Note: The numbering codes here (i, ii, etc) do not correspond with those in the paper, though the items are presented in approximately the same order.

(i) Series #72-80: row 1980. All Texas-Mexico principal components have same 1980 value. Data

 

72

73

74

75

76

77

78

79

80

1977

0.027386

-0.115815

0.029960

0.013702

0.037826

0.003275

0.071702

0.037296

-0.101952

1978

0.092490

-0.001251

0.086672

0.076595

0.022001

0.046141

0.032235

0.024642

0.027261

1979

-0.010550

-0.172530

-0.009996

-0.040788

0.091444

-0.006089

-0.005084

-0.035374

-0.084083

1980

0.023030

0.023030

0.023030

0.023030

0.023030

0.023030

0.023030

0.023030

0.023030

But notice that series #73-80 all start a year earlier than intended:

·        Series #73: starts at 1499  Data

1496

NaN

1497

NaN

1498

NaN

1499

0.122450

1500

-0.057655

1501

-0.069315

1502

0.007860

1503

0.057368

·        series #74-75 starts at 1599 start Data

 

74

75

1596

NaN

NaN

1597

NaN

NaN

1598

NaN

NaN

1599

0.059760

0.103280

1600

-0.047110

-0.046794

1601

0.037995

0.023908

1602

0.022644

-0.017310

1603

0.060597

0.002655

·        series #76-80 start at year 1699 Data

 

76

77

78

79

80

1696

NaN

NaN

NaN

NaN

NaN

1697

NaN

NaN

NaN

NaN

NaN

1698

NaN

NaN

NaN

NaN

NaN

1699

0.070371

0.014897

-0.026822

-0.039449

0.048391

1700

0.079641

0.004216

0.020564

-0.081576

0.128570

1701

-0.079265

0.084556

-0.067874

-0.032629

0.259431

1702

0.091855

-0.023979

0.055814

0.068281

0.138032

1703

0.009087

-0.013712

-0.070385

0.005494

0.130419

Remedy: Columns #73-80 were evidently pasted in at the wrong place and should be shifted down by one cell.

  

(ii) Series #81- 83: row 1980. All Vaganov principal components have same 1980 value. Data

 

81

81

83

1977

0.011704

-0.093460

-0.015599

1978

0.057594

-0.021077

-0.133697

1979

-0.111000

-0.073453

0.024388

1980

-0.040635

-0.040635

-0.040635

But notice that series #81-83 all begin a year earlier than intended:

·        series #81 starts at year 1449 Data

1446

NaN

1447

NaN

1448

NaN

1449

0.049512

1450

-0.004931

1451

-0.009546

1452

-0.058131

1453

0.027849

·        series #82 starts at year 1599 Data

1596

NaN

1597

NaN

1598

NaN

1599

0.046643

1600

-0.085674

1601

-0.049189

1602

-0.006774

1603

0.021100

·        series #83 year 1749 start Data

1746

NaN

1747

NaN

1748

NaN

1749

-0.022483

1750

-0.007668

1751

-0.015818

1752

-0.011613

1753

0.062729

Remedy: Columns #81-83 were evidently pasted in at the wrong place and should be shifted down one row.

  

(iii) Series #84 and #90-92: row 1980. Four ITRDB US principal components have same 1980 value. Data

 

84

85

86

87

88

89

90

91

92

1977

0.021540

0.042500

0.045543

0.063746

-0.082640

0.068311

-0.107215

0.133040

-0.008428

1978

0.022417

0.009168

0.001067

0.024653

-0.037838

0.033733

-0.119638

0.126159

-0.013465

1979

0.048989

0.055037

0.078814

0.065524

-0.063758

0.050831

-0.136815

0.168663

0.028012

1980

0.043453

-0.002933

0.083802

0.049329

0.050349

0.084564

0.043453

0.043453

0.043453

But notice that series #90-92 all begin a year earlier than intended:

·        series #90 starts at year 1599 Data

1596

NaN

1597

NaN

1598

NaN

1599

-0.015548

1600

0.095913

1601

-0.037929

1602

-0.073005

1603

-0.008133

·        series #91-92 year start at year 1749 Data

 

91

92

1746

NaN

NaN

1747

NaN

NaN

1748

NaN

NaN

1749

-0.040770

-0.030561

1750

0.007454

-0.022909

1751

-0.016055

0.074663

1752

-0.075458

-0.045264

1753

-0.054991

0.064213

Remedy: Columns #90-92 were evidently pasted in at the wrong place and should be shifted down one row.

 

 (iv) series 86-89 year start at year 1499 Data 

 

86

87

88

89

1496

NaN

NaN

NaN

NaN

1497

NaN

NaN

NaN

NaN

1498

NaN

NaN

NaN

NaN

1499

-0.017636

0.027544

-0.062917

-0.020833

1500

0.010490

0.035495

0.032561

0.002750

1501

-0.027595

0.013481

0.058521

0.030844

1502

0.029771

0.042843

-0.001940

0.032210

1503

0.005261

0.002590

-0.012206

0.036780

Remedy: Columns #86-89 were likely pasted in at the wrong place and should be shifted down one row.

 

(v) FILLS: Extensive but inconsistent use of extrapolated or interpolated data to cover gaps.

·        series #3, year 1907-1909; data not in underlying source. See further notes below at (2-n).  Data

1905

1.523820

1906

0.569273

1907

-0.042546

1908

-0.044364

1909

-0.046636

1910

-0.839818

1911

0.342000

1912

-0.248909

·        series #3 year 1953-1964 fills. Notice data from 1954 to 1965 rise by 0.001818 each year (except at 1959, by 0.002272). Also 3 fills in 1962-64 overwrite available source data. See further notes below at (2-n). Data

1950

-2.294360

1951

-0.385273

1952

-0.294364

1953

0.796545

1954

-0.398909

1955

-0.397091

1956

-0.395273

1957

-0.393455

1958

-0.391636

1959

-0.389364

1960

-0.387545

1961

-0.385727

1962

-0.383909

1963

-0.382091

1964

-0.380273

1965

0.614727

1966

-0.794364

1967

-1.203450

 

·        Series #6, year 1980 fill. See further notes below at (2-o). Data

1977

-4.990000

1978

-4.805000

1979

-4.717500

1980

-4.717500

1981

NaN

1982

NaN

 

·        Series #45, year 1979-1982 fills Data

1975

14.900000

1976

14.700000

1977

16.100000

1978

15.100000

1979

15.100000

1980

15.100000

1981

15.100000

1982

15.100000

 ·        Series #46, year 1975-1980 fills Data

1971

-1.100000

1972

0.170000

1973

0.640000

1974

-0.430000

1975

-0.430000

1976

-0.430000

1977

-0.430000

1978

-0.430000

1979

-0.430000

1980

-0.430000

1981

NaN

1982

NaN

 ·        Series #50, year 1962-1982  copied from series #49 values in adjacent column Data

 

49

50

1958

0.38000000

0.34000000

1959

-0.15000000

0.45000000

1960

0.280000000

0.020000000

1961

0.120000000

0.550000010

1962

-0.039999999

-0.039999999

1963

0.600000020

0.600000020

1964

-0.779999970

-0.779999970

1965

-0.800000010

-0.800000010

1966

0.289999990

0.289999990

1967

-0.230000000

-0.230000000

1968

-0.949999990

-0.949999990

1969

0.910000030

0.910000030

1970

0.319999990

0.319999990

1971

0.110000000

0.110000000

1972

-0.020000000

-0.020000000

1973

-0.010000000

-0.010000000

1974

-0.079999998

-0.079999998

1975

-0.680000010

-0.680000010

1976

-0.090000004

-0.090000004

1977

0.150000010

0.150000010

1978

-0.140000000

-0.140000000

1979

0.020000000

0.020000000

1980

-0.239999990

-0.239999990

1981

-0.010000000

-0.010000000

1982

0.059999999

0.059999999

 

·        Series #51, year 1977-1980 fills. See further notes below at (2-f). Data

·        Series #52, year 1974-1980 fills. See further notes below at (2-g). Data

·        Series #54, year 1975-1980 fills. See further notes below at (2-h). Data

·        Series #55, year 1979-1980 fills. See further notes below at (2-i). Data

·        Series #56, 1975-1980 fills. See further notes below at (2-j). Data

·        Series #58, 1977-1980 fills. See further notes below at (2-k). Data

·        Also, Series #53, year 1400-1404 fills Data (This one is key, as it lets the Gaspé proxy sneak into the 1400+ group, where it is influential)

 

51

52

53

54

55

56

58

1400

 

 

0.723000

 

 

 

 

1401

 

 

0.723000

 

 

 

 

1402

 

 

0.723000

 

 

 

 

1403

 

 

0.723000

 

 

 

 

1404

 

 

0.723000

 

 

 

 

1405

 

 

0.874000

 

 

 

 

 

 

 

 

 

 

 

 

1971

1.409000

0.460000

1.554000

1.412000

1.048000

1.422000

1.303000

1972

1.257000

0.834000

1.463000

1.388000

1.047000

1.222000

1.388000

1973

1.107000

0.562000

1.618000

1.197000

0.823000

1.071000

1.460000

1974

1.133000

1.104000

1.483000

1.144000

0.949000

1.135000

1.629000

1975

0.932000

1.104000

1.743000

1.366000

1.036000

1.224000

1.613000

1976

1.161000

1.104000

1.577000

1.366000

0.888000

1.224000

1.176000

1977

1.585000

1.104000

1.583000

1.366000

1.119000

1.224000

1.573000

1978

1.585000

1.104000

1.851000

1.366000

1.047000

1.224000

1.573000

1979

1.585000

1.104000

1.618000

1.366000

0.797000

1.224000

1.573000

1980

1.585000

1.104000

2.204000

1.366000

0.797000

1.224000

1.573000

1981

NaN

NaN

1.342000

NaN

NaN

NaN

NaN

1982

NaN

NaN

1.823000

NaN

NaN

NaN

NaN

·        Series #93-99, 1976-1980 fills Data

 

93

94

95

96

97

98

99

1970

-0.05519220

0.03191820

-0.00840994

0.07964660

-0.03334510

0.00628749

0.03972250

1971

0.04456720

0.02654390

0.06869810

0.05834090

0.03159730

0.01390980

-0.00292208

1972

-0.03087120

0.03992700

0.00302668

0.15582700

0.07014980

0.03358720

-0.07598700

1973

-0.02466770

0.11485700

-0.05301170

0.18438500

0.04514380

-0.04919020

-0.05758820

1974

0.03531060

0.07091270

0.00376018

0.11299600

0.01402680

-0.00682486

-0.10635300

1975

0.04918980

0.07842340

-0.02821910

0.16178501

0.02186560

0.02133480

0.00791537

1976

0.04792530

0.07830090

-0.02856930

0.16103400

0.08604140

0.05941720

-0.05302600

1977

0.04792530

0.07830090

-0.02856930

0.16103400

0.08604140

0.05941720

-0.05302600

1978

0.04792530

0.07830090

-0.02856930

0.16103400

0.08604140

0.05941720

-0.05302600

1979

0.04792530

0.07830090

-0.02856930

0.16103400

0.08604140

0.05941720

-0.05302600

1980

0.04792530

0.07830090

-0.02856930

0.16103400

0.08604140

0.05941720

-0.05302600

 But notice that some series were left blank in the late 1970s:

·        Series #102, 1975-1980 missing data Data

·        Series #103, 1975-1980 missing data Data

·        Series #104, 1974-1980 missing data Data

·        Series #106, 1972-1980 missing data Data

 

102

103

104

106

1970

1.161000

1.125000

1.005000

1.638000

1971

1.191000

1.075000

0.946000

0.794000

1972

0.938000

0.833000

0.685000

NaN

1973

0.900000

0.842000

0.763000

NaN

1974

1.197000

1.063000

NaN

NaN

1975

NaN

NaN

NaN

NaN

1976

NaN

NaN

NaN

NaN

1977

NaN

NaN

NaN

NaN

1978

NaN

NaN

NaN

NaN

1979

NaN

NaN

NaN

NaN

1980

NaN

NaN

NaN

NaN

·        Series #112, 1973-1980 missing data Data

1970

1.125000

1971

1.042000

1972

1.273000

1973

NaN

1974

NaN

1975

NaN

1976

NaN

1977

NaN

1978

NaN

1979

NaN

1980

NaN

·        Series #11, 1980 missing data Data

1977

-0.670000

1978

-1.000000

1979

0.330000

1980

NaN

1981

NaN

Remedy: There is no need for such extensive use of artificial data, and moreover fills are applied inconsistently. It might be asserted that they are sparse and will not have much effect on the final result. If so there can be no objection to removing them. Alternatively if they do drive the results this is even more problematic. Either way, they should be removed.

 

 2. Truncations and updates which can be verified at WDCP

The audit trail here is set up as a series of short scripts in R, which read the corresponding data and output (usually a correlation) and referrable here as hyperlinks to series-by-series annotation. Users of R (after loading the MBH98 proxy table using the command below) can simply copy the script into R and the correlation or other index should result. Tweaks for users of Matlab should be apparent to such users.  The location of FTP sources of the various MBH98 series has been by trial-and-error as MBH98 provides no such disclosure. Several locations were identified after this article went to press and are referred to in a postscript here. 

All FTP sources (except the Central England series from Hadley Centre and a Briffa site discussed in the postscript) are from the World Data Center for Paleoclimatology (WDCP), which maintains an excellent collection. WDCP is also sometimes called NGDC—the National Geophysical Data Center. Bruce Bauer, the manager of the paleo program, has been unfailingly co-operative to even the most minute inquiry. While enough FTP sources have been located to make this section of interest, 57 of 112 series have not yet been identified in FTP sources. The major contributor is 28 principal component series calculated by Mann, Bradley and Hughes and never published either digitally or in print. The unavailable data is discussed here.  Surprisingly, some of the more vociferous advocates of aggressive public policy (Hughes, Thompson) have failed to archive their data with WDCP.

For each series in which a source was identified, correlations between the digital series and the MBH series were calculated; the start and finish of each series were examined for truncations or additions and the series were plotted. The scripts hyperlinked below have been condensed to show the material point referred to.  A complete list of FTP references (updating the Appendix) is here.

 a) use of summer data in series #10 (Central England). This is shown first through correlation of >0.99 with JJA series and only 0.62 with annual data and secondly through direct inspection of the 3 series taken together. Graph  Script  Data URL

b) truncation of data from 1659 to 1730 in  series #10 (Central England). This is shown through examination of the two series together. The cold temperatures so deleted are shown by plotting series together. Graph  Script  Data URL

c) high MBH98 data in the 1980s and especially 1987 in series #10 (Central England). This is shown through direct inspection. Graph  Script  Data URL 

d) use of summer data in series #11 (Central Europe). Graph  Script  Data URL

e) truncation of data from 1525 to 1550 in series #11 (Central Europe). This is shown through examination of the two series together. The high early temperatures so deleted are shown by plotting series together. Graph  Script  Data URL

The WDCP identifications of MBH98 series #51-61 (Jacoby northern treeline series) are not shown in MBH98. These identifications are straightforward as shown here. I have downloaded all the WDCP data (decadal format) and converted to R-time series for easier data handling. I've tried to annotate below to show the main issues without requiring this overhead. 

f) MBH98 data for  series #51 (Four Twelve AK) has correlation of 0.86 with WDCP. Comparison of end values shows that WDCP continues to 1990, as compared to MBH end in 1976 (with fills to 1980). Plotting shows that MBH98 has pervasive and increasing over-statement in 20th century values and peaks in the 1920s.  Graph  Script  Data URL

g) MBH98 data for  series #52 (Fort Chimo PQ) has correlation of 0.93 with WDCP. Comparison of end values shows that WDCP continues to 1990, as compared to MBH end in 1976 (with fills to 1980). Plotting shows that MBH98 has pervasive and increasing over-statement in 20th century values. Series peaks in 1960s. Graph  Script  Data URL

h) MBH98 data for  series #54 (Arrigetch AK) has correlation of 0.96 with WDCP. Comparison of end values shows that WDCP continues to 1990, as compared to MBH end in 1976 (with fills to 1980). Plot shows series peak in early 1980s with downturn to series end in 1990. Graph  Script  Data URL

i) MBH98 data for  series #55 (Sheenjek River AK) has correlation of 0.70 with WDCP. Comparison of end values shows that both WDCP and unfilled MBH98 end in 1979. Comparison of start values (and plot) shows WDCP starts much earlier. Considerable overstatement of values in MBH98 in the 1940s and in the 18th century. Graph  Script  Data URL

j) MBH98 data for  series #56 (Twisted Tree, Heartrot Hill (TTHH), Canada has correlation of 0.699 with WDCP. Comparison of end values shows that WDCP continues to 1990, while unfilled MBH98 ends in 1976. Comparison of start values (and plot) shows WDCP starts much earlier. WDCP values peak in the 1960s and reduce sharply thereafter. Increasing MBH overstatement in the 20th century. Graph  Script  Data URL

k) MBH98 data for  series #58 (Coppermine River, Canada has correlation of 0.99 with WDCP. MBH fill three years (1978-1980), but otherwise coverage period is the same. Values nearly identical at beginning but pervasive changes later in the series.  Graph  Script  Data URL

l) MBH98 data for series #1 Burdekin River, Australia coral fluorescence has correlation of 0.42 with WDCP series. Lough (pers. comm. Oct. 2003) confirms validity of WDCP series over earlier data. Plot shows visual coherence, but considerable shifting.  Graph  Script  Data URL

m) MBH98 data for  series #2 (Great Barrier Reef) is coral calcification, not coral thickness (Lough, pers. comm., Oct. 2003). There is a correlation of 0.99 between series #2 and the average calcification at WDCP of the following 5 corals for the period 1615-1982: Abraham Reef, Britomart Reef, Havannah Island, Lodestone Reef and Sanctuary Reef. The MBH data seems to be Z-transformed, although the basis of the Z-transform is not clear. Graph  Script  Data URL: Abraham Reef     Britomart Reef     Havannah Island     Lodestone Reef     Sanctuary Reef

n) MBH98 data for  series #3 Urvina Bay, Galapagos coral δO18 has correlation of -0.9992951 with WDCP series - which is reversed in sign during transformation. MBH overwrite actual data in 1962-64 and fill for 1907-1909 and 1953-61 as noted above.  The missing data results from a splice between two corals, which are spliced by adjusting the readings of the second coral. Graph  Script  Data URL

o) MBH98 data for  series #6 Vanuatu coral δO18 has correlation of 0.93 with WDCP series. MBH have one filled year in 1980. Graph  Script  Data URL

p) MBH98 data for series #7, New Caledonia  δO18 has correlation of 0.618 with WDCP data. Graph  Script  Data URL

q) MBH98 data for series #8, Secas, Panama  δO18 has correlation of 0.983 with WDCP annualized series (annual data calculated from WDCP 10 per year data). Graph  Script  Data URL

r) MBH98 data for series #9, Secas, Panama  δC13 has correlation of 0.991 with WDCP annualized series (annual data calculated from WDCP 10 per year data). Graph  Script  Data URL

s) MBH98 data for series #21, grid-box 42.5N, 92.5W has correlation of 0.889 with JB92 Minnesota (adjacent grid box) annual data. There are many differences in the plotted series. Graph  Script  Data URL

t) MBH98 data for series #23, grid-box 47.5N, 7.5E has correlation of 0.81 with JB92 Geneva annual data, which has identical start date (1753) and location. There are many differences in the plotted series including a notable downspike in the MBH data in early 19th century not present in JB92 data. Graph  Script  Data URL

u) MBH series #26, grid-box 52.5N, 17.5E has no counterpart location in JB92 Table 13.1.

v) MBH98 data for series #27, grid-box 57.5N, 17.5E has correlation of >0.99 with JB92 Stockholm annual data, which has identical start date (1756) and location. The MBH series is linearly transformed from the JB92 series. Graph  Script  Data URL

w) MBH98 data for series #28, grid-box 57.5N, 37.5E has correlation of 0.96 with JB92 Leningrad annual data, which has identical start date (1752) and location. The MBH series is transformed from the JB92 series.  Graph  Script  Data URL

x) MBH series #29, grid-box 62.5N, 7.5E has no counterpart location in JB92 Table 13.1.

y) MBH98 data for series #30, grid-box 62.5N, 12.5E has correlation of 0.998 with JB92 Trondheim annual data, which has identical start date (1761) and location. The MBH series is transformed from the JB92 series. Graph  Script  Data URL

z) JB92 series for Central England , Berlin , Sverdlovsk and Toronto (all digitally available at WDCP) are compared to MBH series #21-31 and no correlations are found to permit identification.

aa) MBH98 data for series #35, grid-box precipitation 42.5N, 2.5E has correlation of 0.95 with JB92 Marseilles (43.3N, 5.4E) annual data, which has identical start date (1749) and is one grid-box to the east. Both the JB92 series at WDCP and MBH series are transformed, but transformations are different. Graph  Script  Data URL

ab) MBH98 data for series #37, precipitation 42.5N, 72.5W has correlation of 0.92 with JB92 Paris annual data, which has identical start date (1770). Both the JB92 series at WDCP and MBH series are transformed, but transformations are different. Graph  Script  Data URL

ac) JB92 series

ad) MBH98 data for series #43, Tasmania T-reconstruction has correlation of 0.82 with updated WDCP series. Plot shows visual coherence, but considerable shifting.  Graph  Script  Data URL

ae) MBH98 data for series #65, Tarvagatny Pass, Mongolia has correlation of 0.94 with updated WDCP series. MBH data shows increasing over-estimate in 20th century.  Graph  Script  Data URL

af) MBH98 data for series #105, INDI008X is an incorrect label for WDCP indi002x.  Correlation is 0.83.  Graph  Script  Data URL

ag) MBH98 data for series #112, SWED002B is WDCP swed002.  Correlation is 0.977.  Graph  Script  Data URL

Series which were successfully located in digital form in the MBH98 form are noted here; comments on digitally unavailable series are here

Remedy: In every case the most updated and complete records from WDCP are used, superceding the corresponding records in the MBH98 data base.  

Postscript:

The following additional obsolete data was identified after the article went to press and is not incorporated into the analysis.

ah) A non-WDCP location for series #65 Tarvagatny Pass, Mongolia was located - series ar at the URL. MBH98 data is truncated for the period 1466 to 1550, but otherwise duplicates Briffa series ar:   URL  As noted in (ae), this series is superceded by the NGDC series, archived June 24, 2002.

ai) A non-WDCP location for series #66 - Yakutia is series br at Briffa. The series shown at Briffa corresponds to Figure 8 of Hughes et al. (1999), which commences in 1400, truncating approximately 150 years of data from Figure 7.  Hughes et al. (1999) reported data that they measured data from 7 sites in the Yakutia area, but have failed to contribute this data to WDCP. 

aj) A non-WDCP location for series #67 - Fennoscandia was located - series e2r (!) at Briffa. (Briffa has reversed the labels on his webpage as at Oct. 20, 2003.  This was verified by comparison to the print articles.) MBH98 data is identical to Briffa f2r in the overlap period.  The updated version of this series  - series e1r (!) has a correlation of only 0.63 with the MBH version. 

ak) A non-WDCP location for series #68  Polar Urals was located - series f2r (!) at Briffa. (Briffa has reversed the labels on his webpage as at Oct. 20, 2003.  This was verified by comparison to the print articles.) MBH98 data is identical to Briffa f2r in the overlap period.  The updated version of this series  - series f1r (!) has a correlation of only 0.38 with the MBH version and of 0.28 with the longer version. 

al) Several other sites were posted at  Briffa and would seem to be candidates for MBH98: Taimir, Athabaska. The reasons for exclusion are not given. However, they are applied in Bradley, Hughes and Diaz (2003).

 

 3. Tree Ring Principal Components

Five separate principal component regions were identified within the MBH98 database: Texas-Oklahoma (#69-71), Texas-Mexico (#72-80), ITRDB US (#84-92), South America (#93-95) and Australia-(New Zealand) (#96-99).

The sites for each region are identified at MBH Supplementary Information.  No sites for Texas-Oklahoma or Texas-Mexico were given there, but WDCP identifications for the sites in these region were easily located and are listed here. All sites in the other three regions, except immaterially one of 232 US sites -AR045, were located at WDCP. Digital site lists are as follows: Texas-Oklahoma, Texas-Mexico, ITRDB US, South America and Australia-(New Zealand) as well as for MBH99 ITRDB US.

The site chronologies (*.crn) data from WDCP was collated into time series for each region, truncating the US data at 1400.  Digital collations are as follows: Texas-Oklahoma, Texas-Mexico, ITRDB US, South America and Australia-(New Zealand).  A collation is also done for the MBH99 ITRDB US data.

Conventional principal component calculations, which MBH98 claim to use, require that there be no missing data.  There is little relationship between the periods in which MBH principal components are calculated and the period during which all selected sites in the region are available as shown here.

Tree ring data is conventionally standardized to a mean of 1000 (with no negative values).  Although it is not disclosed by MBH, they carry out a Z-transformation on the collated data. This is established both by the range of values and by a very close replication of the MBH99 principal components.  MBH99 PC1  MBH99 PC2 MBH99 PC3.

  Accordingly, prior to carrying out a principal component calculation, the collated data is scaled.  A principal components analysis is carried out for each region and the same number of principal components collected as in the MBH98 collection.  The explained variance is calculated.  MBH do not disclose the eigenvectors corresponding to their principal components. Given the MBH98 PCs, the eigenvectors which maximize explained variance are calculated; the explained variance using the MBH PCs and these calculated eigenvectors is then calculated.  The script to carry out the calculations in this section is here.  Graphics comparing the MBH and recalculated PCs are here.

Texas-Oklahoma PC1  PC2 PC3
Texas-Mexico PC1  PC2 PC3  PC4 PC5  PC6 PC7  PC8 PC9
ITRDB US PC1  PC2 PC3  PC4 PC5  PC6 PC7  PC8 PC9
South America PC1  PC2 PC3
Australia PC1  PC2 PC3  PC4

A summary of explained variance is here.  A summary of correlations is here.

Remedy: The principal component calculations above yield higher explained variance levels in every case, and hence are used instead of the MBH98 PCs.

END RESULT: The data set incorporating all the above remedies is HERE.

<< Return to Main page.