Friday, January 29, 2016

Coursera Regression Course - Assignment 2

A major life depression is associated with
a higher daily consumption of alcohol

All data from the NESARC Study

Categorical Explanatory Variable

We have a categorical variable encoded with No=0, implying that subjects with no major depression in life are at the baseline 0.
3549-3549
MAJORDEPLIFE
MAJOR DEPRESSION - LIFETIME (NON-HIERARCHICAL)
----------------------------------------------
35254 0. No
7839 1. Yes
----------------------------------------------

Quantitative Response Variable

The amount of ethanol (alcohol) consumed averagely in the past year is our response variable.
3675-3682
ETOTLCA2
AVERAGE DAILY VOLUME OF ETHANOL CONSUMED IN PAST YEAR, FROM ALL TYPES OF ALCOHOLIC BEVERAGES COMBINED
(NOTE: Users may wish to exclude outliers)
--------------------------------------------------
0.0003 - 219.9555 Ounces of ethanol
Blank Unknown
--------------------------------------------------

We see that there is a large number of abstinent people in the dataset. We still don't know if they are associated with depression or not. It would be interesting to see if they are more the one or the other.

Linear Regression Test

                            OLS Regression Results                            
==============================================================================
Dep. Variable:               ETOTLCA2   R-squared:                       0.001
Model:                            OLS   Adj. R-squared:                  0.001
Method:                 Least Squares   F-statistic:                     18.86
Date:                Fri, 29 Jan 2016   Prob (F-statistic):           1.42e-05
Time:                        15:57:58   Log-Likelihood:                -59022.
No. Observations:               26655   AIC:                         1.180e+05
Df Residuals:                   26653   BIC:                         1.181e+05
Df Model:                           1                                         
Covariance Type:            nonrobust                                         
================================================================================
                   coef    std err          t      P>|t|      [95.0% Conf. Int.]
--------------------------------------------------------------------------------
Intercept        0.5366      0.015     35.377      0.000         0.507     0.566
MAJORDEPLIFE     0.1474      0.034      4.342      0.000         0.081     0.214
==============================================================================
Omnibus:                    84158.156   Durbin-Watson:                   2.010
Prob(Omnibus):                  0.000   Jarque-Bera (JB):      20279387185.774
Skew:                          50.216   Prob(JB):                         0.00
Kurtosis:                    4274.926   Cond. No.                         2.62
==============================================================================

The R-squared number is rather low, sugesting that only 0,1% of the "Etanolic Intake" can be explained by a "major depression in life". At least the F-statistic is so low that we can assume the small effect is statistically significant. 
Now I would not go further but lets assume that R-squared is higher, in that case we believe that a major depression in a lifetime can predict Beta=14.74% increase in alcohol intake, with a high statistical significance of P>0.000

Graph: a major life depression could increase average ethanol intake from 54% to 68%

Program Code and Output

In [1]:
%matplotlib inline

import numpy
import pandas
import statsmodels.api as sm
import seaborn
import matplotlib.pyplot as plt
import statsmodels.formula.api as smf 
In [2]:
# bug fix for display formats to avoid run time errors
pandas.set_option('display.float_format', lambda x:'%.2f'%x)
In [3]:
data = pandas.read_csv('nesarc_pds.csv', low_memory=False)
In [4]:
#setting variables you will be working with to numeric
data['p'] = pandas.to_numeric(data['MAJORDEPLIFE'], errors='coerce')
In [5]:
data['ETOTLCA2'] = pandas.to_numeric(data['ETOTLCA2'], errors='coerce')
In [6]:
subGroupEthanol = list(filter(lambda x: x > 10.0, data['ETOTLCA2'].dropna()))
seaborn.distplot(subGroupEthanol)
Out[6]:
In [7]:
reg1 = smf.ols('ETOTLCA2 ~ MAJORDEPLIFE', data=data).fit()
print (reg1.summary())
                            OLS Regression Results                            
==============================================================================
Dep. Variable:               ETOTLCA2   R-squared:                       0.001
Model:                            OLS   Adj. R-squared:                  0.001
Method:                 Least Squares   F-statistic:                     18.86
Date:                Fri, 29 Jan 2016   Prob (F-statistic):           1.42e-05
Time:                        17:11:29   Log-Likelihood:                -59022.
No. Observations:               26655   AIC:                         1.180e+05
Df Residuals:                   26653   BIC:                         1.181e+05
Df Model:                           1                                         
Covariance Type:            nonrobust                                         
================================================================================
                   coef    std err          t      P>|t|      [95.0% Conf. Int.]
--------------------------------------------------------------------------------
Intercept        0.5366      0.015     35.377      0.000         0.507     0.566
MAJORDEPLIFE     0.1474      0.034      4.342      0.000         0.081     0.214
==============================================================================
Omnibus:                    84158.156   Durbin-Watson:                   2.010
Prob(Omnibus):                  0.000   Jarque-Bera (JB):      20279387185.774
Skew:                          50.216   Prob(JB):                         0.00
Kurtosis:                    4274.926   Cond. No.                         2.62
==============================================================================

Warnings:
[1] Standard Errors assume that the covariance matrix of the errors is correctly specified.
In [8]:
# listwise deletion for calculating means for regression model observations
sub1 = data[['ETOTLCA2', 'MAJORDEPLIFE']].dropna()

# group means & sd
print ("Mean")
ds1 = sub1.groupby('MAJORDEPLIFE').mean()
print (ds1)
print ("Standard deviation")
ds2 = sub1.groupby('MAJORDEPLIFE').std()
print (ds2)

# bivariate bar graph
seaborn.factorplot(x="MAJORDEPLIFE", y="ETOTLCA2", data=sub1, kind="bar", ci=None)
plt.xlabel('Major Life Depression')
plt.ylabel('Mean Number AVERAGE DAILY VOLUME OF ETHANOL CONSUMED IN PAST YEAR')
Mean
              ETOTLCA2
MAJORDEPLIFE          
0                 0.54
1                 0.68
Standard deviation
              ETOTLCA2
MAJORDEPLIFE          
0                 1.44
1                 4.03
Out[8]:
In [ ]:
 

Wednesday, January 20, 2016

First Assignment Coursera - Regression Models in Practice


I have chosen the National Epidemiologic Survey on Alcohol and Related Conditions (NESARC)

http://pubs.niaaa.nih.gov/publications/AA70/AA70.htm


1. Describe the sample

a) Describe the study population (who or what was studied).

NESARC contains an extensive battery of questions about present and past alcohol consumption, AUDs, and the use of alcohol treatment services. NESARC also included similar questions related to tobacco and illicit drug use (including nicotine dependence and drug use disorders) as well as questions designed to determine a wide variety of psychiatric disorders such as major depression, anxiety disorders, and personality disorders. 
 Furthermore, NESARC contained questions that operationalized the criteria set forth in the American Psychiatric Association’s Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition (DSM–IV) for the following psychiatric disorders:
  • Five mood disorders (major depressive disorder, bipolar I and bipolar II disorders, dysthymia, and hypomania)
  • Four anxiety disorders (panic with and without agoraphobia, social phobia, specific phobia, and generalized anxiety)
  • Seven personality disorders (avoidant, dependent, obsessive–compulsive, paranoid, schizoid, histrionic, and antisocial disorders).

b) Report the level of analysis studied (individual, group, or aggregate).

Individuals were randomly selected from a systematic sample of group quarters in these PSUs.The Census 2000 Group Quarters Inventory formed the sampling frame for the group quarters portion of the NESARC sample. 

c) Report the number of observations in the data set.

43093 Total, 8209 1. Northeast 8991 2. Midwest 16156 3. South 9737 4. West
d) Describe your data analytic sample (the sample you are using for your analyses).
397-398 S2AQ8A HOW OFTEN DRANK ANY ALCOHOL IN LAST 12 MONTHS
 --------------------------------------------- 
 1865 1. Every day 
 1210 2. Nearly every day 
 2619 3. 3 to 4 times a week 
 2914 4. 2 times a week 
 3261 5. Once a week 
 3557 6. 2 to 3 times a month 
 2663 7. Once a month 1805 8. 7 to 11 times in the last year 
 3210 9. 3 to 6 times in the last year 
 3637 10. 1 or 2 times in the last year 
 205 99. Unknown 
16147 BL. NA, former drinker or lifetime abstainer 
---------------------------------------------------------------------------------------

2. Describe the data collection procedure

a) Report the study design that generated that data (for example: data reporting, surveys, observation, experiment).

Surveys, Data were collected in face-to-face, computer-assisted personal interviews conducted in respondents’ homes. The NESARC response rate was 81 percent.

b) Describe the original purpose of the data collection.

The major purposes of the Wave 1 and Wave 2 NESARC are to:
  • Determine the prevalence, incidence, stability, and recurrence of AUDs and their associated disabilities in the general U.S. population.
  • Estimate the magnitude of health disparities in AUDs and their associated disabilities among population subgroups defined by gender, race/ethnicity, disability, sexual orientation, age, and socioeconomic status.
  • Estimate the size, characteristics, and changing nature of populations of special concern, including alcohol abusers and other people in the general population who are impaired or affected by the use of alcohol (e.g., those engaging in binge drinking or impaired driving).
  • Estimate changes in AUDs and their associated disabilities over time, and identify factors associated with the natural history of AUDs.
  • Determine the number of people receiving alcohol treatment through various treatment programs and services, including those not otherwise represented in surveys of treatment facilities; measure the unmet need for alcohol treatment services; and identify barriers to seeking treatment.
  • Determine the associations between AUDs and their major physical and mental disabilities, differentiating drug-induced disorders from those reflecting true, independent mental conditions.
  • Determine the boundaries between safe and hazardous drinking levels and patterns for various types of AUDs and their associated medical, social, and psychological sequelae.

c) Describe how the data were collected.

NESARC participants are variable in lifestyle and ages.

To ensure that minority and special populations were well represented in the sample, NESARC oversampled Blacks, Hispanics, and young adults ages 18–24. As a result, the survey produced enough minority respondents to answer questions of race/ethnic disparities in comorbidity and access to health care services

d) Report when the data were collected.

In 2001—2002, the National Institute on Alcohol Abuse and Alcoholism (NIAAA) conducted the first wave of the National Epidemiologic Survey on Alcohol and Related Conditions (NESARC),The second wave of the survey took place from 2004 to 2005. 

e) Report where the data were collected.

They represented all regions of the United States. In addition to sampling people living in traditional households, NESARC investigators questioned military personnel living off base and people living in a variety of group accommodations such as boarding or rooming houses and college quarters. By including these different types of housing, the investigators were able to obtain data on people not typically captured by household surveys.

 3. Measures section describing your variables and how you managed them to address your own research question

a) Describe what your explanatory and response variables measured.

explanatory variable:

397-398 S2AQ8A HOW OFTEN DRANK ANY ALCOHOL IN LAST 12 MONTHS
 --------------------------------------------- 
 1865 1. Every day 
 1210 2. Nearly every day 
 2619 3. 3 to 4 times a week 
 2914 4. 2 times a week 
 3261 5. Once a week 
 3557 6. 2 to 3 times a month 
 2663 7. Once a month 1805 8. 7 to 11 times in the last year 
 3210 9. 3 to 6 times in the last year 
 3637 10. 1 or 2 times in the last year 
 205 99. Unknown 
16147 BL. NA, former drinker or lifetime abstainer 
---------------------------------------------------------------------------------------
b) Describe the response scales for your explanatory and response variables.
c) Describe how you managed your explanatory and response variables.

response variable

 SECTION 9: GENERALIZED ANXIETY (GENERAL ANXIETY) -
 3020-3020 S9Q1A EVER HAD 6+ MONTH PERIOD FELT TENSE/NERVOUS/WORRIED MOST OF TIME
 3128 1. Yes 38225 2. No 1740 9. Unknown

 3021-3021 S9Q1B EVER HAD 6+ MONTH PERIOD FELT VERY TENSE/NERVOUS/WORRIED MOST OF TIME ABOUT EVERYDAY PROBLEMS
 358 1. Yes 37867 2. No 1740 9. Unknown 3128 BL. NA, had 6+ month period feeling tense/nervous/worried most of the time

 3022-3022 S9Q31 IN WORST PERIOD, EVER WORRY A LOT ABOUT THINGS YOU USUALLY DIDN'T WORRY ABOUT
2336 1. Yes 1115 2. No 35 9. Unknown 39607 BL. NA, never or unkn. if ever had 6+ month pd. being tense/nervous/worried

 3023-3023 S9Q32 IN WORST PERIOD, EVER WORRY ABOUT MORE THAN ONE THING
 2644 1. Yes 808 2. No 34 9. Unknown 39607 BL. NA, never or unkn. if ever had 6+ month pd. being tense/nervous/worried

 3024-3024 S9Q33 IN WORST PERIOD, EVER FIND IT DIFFICULT TO STOP BEING TENSE/NERVOUS/WORRIED
 2715 1. Yes 742 2. No 29 9. Unknown 39607 BL. NA, never or unkn. if ever had 6+ month pd. being tense/nervous/worried

 3025-3025 S9Q34 IN WORST PERIOD, EVER WORRY ABOUT THINGS THAT WERE VERY UNLIKELY TO HAPPEN
1418 1. Yes 2019 2. No 49 9. Unknown 39607 BL. NA, never or unkn. if ever had 6+ month pd. being tense/nervous/worried

c) Describe how you managed your explanatory and response variables.
confounders:

I will control for different confounders like:

 Page 302: SECTION 4A: MAJOR DEPRESSION (LOW MOOD I)
Page 316: SECTION 4B: FAMILY HISTORY (III) OF MAJOR DEPRESSION
Page 319: SECTION 4C: DYSTHYMIA (LOW MOOD II)
Page 329: SECTION 5: MANIA OR HYPOMANIA (HIGH MOOD)
Page 341: SECTION 6: PANIC DISORDERS AND AGORAPHOBIA (ANXIETY)
Page 355: SECTION 7: SOCIAL PHOBIA (SOCIAL SITUATIONS)
Page 371: SECTION 8: SPECIFIC PHOBIA (SPECIFIC SITUATIONS)
Page 386: SECTION 9: GENERALIZED ANXIETY (GENERAL ANXIETY)
Page 403: SECTION 10: PERSONALITY DISORDERS (USUAL FEELINGS/ACTIONS)
Page 419: SECTION 11A: ANTISOCIAL PERSONALITY DISORDER (BEHAVIOR)

Sunday, September 1, 2013

Which Travel Guide for Tyrol? To Germans does not matter

My endevor to build an tourism indicator with online booksales led to a first interesting result.
In the category of worldwide travelguides, you see that the rank number (more is worse) increases comparably fast.
Whereas within the Tyrol category there are alot of books doing comparably well. The rank number increases slow.

World travel guides 
Italy travel guides
South Tyrol travel guides
Tyrol travel guides

Monday, January 9, 2012

See you at G+

As long as I have not a longer article to publish you find daily rants and ramblings on certainty, uncertainty and ignorance under https://plus.google.com/u/0/103983654382933904290/posts

Saturday, August 27, 2011

Easy way to stop smoking, this time it works!

It was Supposed to be Easy
Like millions I read the famous book by Allen Carr "Easy way to stop smoking". Like millions, I tought: he has a point. Many quited, but I went on. Now I'v found a secret weapon of mass responsibility, where I can change my behavior for good. The weapon is called Commitment Contract.

UPDATE: it's now over a month since I am without smoking. It was tough, but you just stick with it.

Why Quit?
The consequences of my behavior are well studied:
Men born in 1900-1930 who smoked only cigarettes and continued smoking died on average about 10 years younger than lifelong non-smokers. Cessation at age 60, 50, 40, or 30 years gained, respectively, about 3, 6, 9, or 10 years of life expectancy. 
 [ study prolonged over a century]


This is the current best knowledge:
  1. Quit before 40 and chances are that your smoking has no consequences. 
  2. Quit before 50 and you shorten your life for -5 years.
  3. Quit before 60 and you shorten your life for -7 years.
  4. Don't quit and you shorten your life for -10 years.
Why Smoking?
Addiction is all about dopamine, a hormone that makes you feel happy. A cigarette gives you a small dose of Nicotine which in turn kickstarts the increase of adrenaline, serotonin and dopamine. You are more aware, more reactive and more happy after smoking. Just a little, not much.
To learn what is good for one-self is crucial for the sophisticated survival machines we are. Therefore a mamals brain learns to seek dopamine rewards. With every cigarette you learn that smoking is good, and the absence of nicotine worries you. The phenomena of withdrawal.
Although the damages it does are way superior to the possible benefits, not everything is bad with smoking: current medicine knowledge is that nicotine helps against alzheimer's and parkinson's. But I also love the nicotine rich tomatoes and potatoes and I started to complete my nutrition with Omega 3 DHA and EPA (makes you smart and protects against the major civilatory deseases like Alzheimer and cardio-vascular deseases).

How to Quit?
Good news: smoking does not provide increasing dopamine returns.
For a binge drinker, every drink augments the desire of another drink.
Even for a chain smoker this is not the case. The craving for a cigarette might not disappear but it certainly does not increase.
Decreasing Returns
The first cigarette is the most wanted.
The addiction develops more as a decrease in the inteval between two smokes.

This is important. We now know the maximum craving is the longing for the first cigarette. We can measure it! What is the unit of measurement? An economist would say: money. I'd say: money is not enough, but a good start. How much would you give me for a cigarette if you really need one?

  1. Would you give me 100 Euros? 
  2. Would you give me 50 Euros? 
  3. Would you give me 10 Euros? 
  4. ...

There is an amount which you would pay when craving hard. I decided mine is somewhere around 5 to 10 Euros. This is a measure of your addiction.


Hedging against Smoking
If you want to stop smoking you have to make sure that for each cigarette you smoke, you will lose more value than your addiction treshold. For example you can  formulate a contract: "I pay 10 Euros for every cigarette I smoke." I did this, but I did more:

  1. I determined a period for the contract: until 2012-01-01 (Short! Somehow I was not ready for more)
  2. I convinced a friend to be my referee, I gave him 100 Euros to take under ward. 
  3. I declared in my contract that I have to give 10 Euros to a detestable political organisation for every cigarette I smoke. This is anti-charity. This is admitting we are not homo economicus - it is is for the social animal in us. Donating the money to an anti-charity pushes the commitment from being only a higher price (I pay for it, so I have the right to smoke it) to something that would weaken your identity if you fail.
My incentives to resist the Nicotine-crave are now bigger than the temptations.
During the first week I felt the urge to smoke strongly. Sometimes I couldn't sleep well, I was nervous. But it really never came to my mind to smoke a cigarette. Its clear that I do not want to donate to the detestable political organisation. The idea of doing that kept me away from the cigarettes. I've externalized my right to smoke by installing a penalty. If it works for the traffic police it's certainly good for keeping me away from bad habits.


The ideas presented here come from Ian Ayres' Carrots and Stick. The idea is genius but the book is a little bit boring, I'd recommend his other book "Supercrunchers".

Thursday, July 28, 2011

Which day I lose weight

I could swear it is weekend's.
But making a 'seasonal' plot, thus aggregating all data for each weekday, showed I do not so good on weekends. Especially Sundays. I am supposed to do sport on sundays. But, hey we had really bad weather the whole July. And some social obligations too...

Is it my day?
Boxes and Wiskers Plot of weekdays: the horizontal line in the middle is the  median. Upper box border is the 75% quartile, lower border the 25% quartile,so that 50% of values are between 25-75%. Lines outside are the the next 2-98 percentiles. Dots are outliers. Forget the legend, 0.8 is only the opacity of the boxes.


So in hindsight my worst day is Sunday and my best is Wednesday. How does it relate to kcal left to base rate? How would a weekday regression look like?
A weak start
The regression tells that sunday and monday are fatty
The bigger the dots, the more calories left to base rate. 

I even cheated one sunday because of overnight stay in the mountains, including booze, I didn't record.
Disclaimer: You don't know what a mountain is, if you where not there. And I know the evidence is 'thin', five data-points per weekday is not what you should base a scientific law upon