Sunday, March 8, 2015

Pulling Stock Data through R/YQL

After i got the list of all tickers of products around the world, I started pulling data and doing some analysis. The easiest of the analysis were
1) Finding the spread
2) Finding the correlated products

I chose stocks from 2 countries
a) BSE Stocks from India
b) German Stocks from the DE Bourse
c) Currency data all over the world

Here are a few interesting points
From the analysis from the year 2013 onwards, the currencies CAD and GTQ had the highest negative correlation. In order to verify it, I pulled data from the website and it matches

On the flip side, the highly correlated currencies were DKK and XAF


One more observation is that most of the OPEC countries have their currency highly correlated to the dollar trying to find a correlation between them and other currencies s similar to finding correlations between USD and other currencies

The R script which i used to find out these is located here and here. The first file is to calculate the spread/deviation and the second is to download the data. The scripts are pretty simple, one needs to change the directories though for their use

If you want to use excel to extract the stock information, then the best method would be YQL .It is easy to use and you have all sorts of information in  a structured format

If you want to pull out historical prices with the following specifciations
Symbol/Ticker : symbol
StartDate
EndDate

then the correpsonding weblink to extract this info would be ( the below is in excel vba format )
"http://query.yahooapis.com/v1/public/yql?q=select%20*%20from%20yahoo.finance.historicaldata%20where%20symbol%20in%20%28%27 " & symbol & "%27%29%20and%20startDate%20=%20%27" & startDate & "%27%20and%20endDate%20=%20%27" & endDate & "%27&diagnostics=true&env=store://datatables.org/alltableswithkeys"

I have created an excel sheet where you can give a list of the stock tickers and the start and end date and it will download the neccessary information within the excel sheet

You can download the sheet here and tweak it as per your requirements. Enjoy

Wednesday, February 25, 2015

Cutting of Trees at Aarey Milk Colony, Mumbai

The recent decision by MMRDA to cut down 3000 trees in the Aarey Colony is a bad one. It seems they have little idea of the importance of the aarey forest land around Mumbai

With a population of 20000000 Mumbai emits around 3 crore tons of carbon per year. Most of it is thrown into the atmosphere. A large part is sequestered by the Sanjay Gandhi National Park. However it is one trees work at a time.

Due to the 3000 trees being cut , 52 tons of carbon will be emitted into the atmosphere. And on an annual basis a 100 tons of carbon which would have been sequestered will now remain in the atmoshphere


Monday, January 12, 2015

Movie Stats .. Bollywood

Hi, I have been recently thinking of an area to analyse and after watching some bad movies, I decided to give a go for bollywood movies. Can we crunch it through numbers/stats to fin out whether the movie will do good or bad ( irrespective of what the movie is about). This project is comprehensive and it will take time to collect data, so i will go slowly

I took the 2001-2010 decade and tried to analyse the move gross by months. We see that that the month of Dec did not hold a lot of promise. It was obviously pre-Aamir Khan era. From 2007 there has been a surge in the gross collected in the month of December. A jump of almost 150% to 100 crore rupees



In the above images we can see a lot of shifts

1) Increase in movie gross Revenue ( in crores )

2) Change in the monthly Gross Revenue ( in crores )

It can be easily seen that purely on the basis of Month, we will not be able to ascertain the earnings. E.g. In the era 2002 to 2009, there were few months that would steal away most of the revenues. In 2007-2008 alone Aug, Oct and Dec stole away 50% of the yearly gross. In 2013 the top 3 months have taken away 38% of the revenue. The figure was even less in the years 2012 and 2011. Movies have begun to spread out evenly and there are more opportunities for newcomers and new genres

On running a simple linear regression between the
a) Earnings of Dec : Dependant Variable
b) Earnings of Aug, Sep, Oct, Nov : Independant Variables


  Coefficients
Intercept 11.28481646
Aug 0.281777793
Sep 0.280503546
Oct -0.094513125
Nov -0.419597559

It is inversely related to the month of November ( highly inverse )

There are more factors at play
1) Director/Producer
2) Actor
3) Festival
4) etc.. etc

I will be taking all of these into consideration in my next article. Please let me know if you have some other suggestions/ideas

Thursday, January 8, 2015

Pulling Company Information through R


For some time now, I was grappling  with getting company fundamental ratios to be able to do analysis. I have written a small R script that downloads the necessary market data for that stock ( in the example AAPL ). The source of the data is YAHOO Finance

After execution the R workspace will contain 2 data frames
ratios    -> Quantitative Data
ratios2  -> Qualitative Data

Download File


I am working on getting more balance sheet related information. It is always better to have more information while doing analysis

If anybody wants to have the complete list of companies enlisted in NSE, you can get it here Download File

After searching around for some time, I wrote a script that will scrape and list the tickers etc for products around the world. The current list contains of 110,000 products. You can download the Excel file here Download File Timestamp : 2nd March,2015 ( I refrained from using csv, because the company name consisted of all sorts of characters which can meddle with string separation :) )

Thursday, December 25, 2014

Cricket Stats

I was playing around with cricket data for some time now, and there are some interesting observations

1) Average Runs scored per match YOY

2) Average Runs scored when tier2 teams won the toss and elected to bat


  3) Average Runs scored when tier2 ( Kenya, Zimbabwe) teams won the toss and elected to bowl



The above stats would state that maybe 2014 was a good year for cricket, but we have forgotten one important measure. The number of matches. More the number of matches, more the normalization ( averaging out )

In the following graphs, I plotted the average runs and the number of matches YOY for all the 3 forms of cricket

1) ODI

 2) T20
3) TEST




We see that for the year 2014 there is a surge in the number of runs scored in all 3 versions of the game. However, the data has fewer matches so statistically we cannot say for sure that 2014 has been a very good year.
However, I always thought that through the years, the number of runs scored have always been increasing, but that is not what we observer, especially between 2009 and 2013. Let us dig deeper


We see that in the year 2014, there were a lot of matches > 30 with the total being above 500. This drove the average up. From the graph we can make some more observations.

Now, we will  be applying similar stats on the following factors
1) Wickets fallen in a match
2) Runs scored till fall of first wicket
3) % of highest wicket partnership vs total runs

Will come up with the details


Potholes at Traffic Signals

Everybody must have at some time or the other faced a situation where we are stuck at a traffic signal, idly looking at the smoke coming out of the nearby truck. We keep looking ahead to notice the reason the traffic is held up. When we reach the signal, we find it was a pothole. How could the pothole hold up the entire traffic? It is very well true

Assumpions

average length of the vehicle 2.5
average speed of vehicles ( km / hr ) 30
time for whch the signal is open ( seconds ) 60
time for whch the signal is closed ( seconds ) 180
idling losses ( average ) mL/hr 500
Road Length ( in m ) 500



The above chart calculates the loss in  litres of petrol, per crossing, per period ( signal open + signal close ) assuming that there are 3 lanes. The x axis is the average speed of  vehicles, and the y axis shows the loss of oil due to "Idling"

The following link will now give you vehicle population of India. A graphical representation of the same is as follows
We see that there has been a humongous increase in 2 wheeler and small 4 wheelers in the last 2 decades. Assuming that out of this 50% vehicles get stuck at traffic signals (which is highly optimistic) the loss of oil runs into millions per year

With the following assumptions

Idling losses ( ml/Hr  ) 200
Average time wasted at traffic/pothole ( hr/day ) 0.5





The total cost currently would stand not less than 100 million INR. It is obviously much more but I tried to arrive at a calculated value with a conservative approach

Wednesday, December 24, 2014

Crime Against Women

Crime against women is on the rise, but what we read in the newspapers about violent attacks in only a tip of the iceberg. People might think molestation/rape is the foremost crime, but it is not. According to the data pulled from data.gov.in

1) DOWRY is one of the foremost reasons of crime against women and Andhra Pradesh has a huge chunk of it.  The numbers include cases registered for dowry prevention act

2) The following graph sums up all the registered cases by State. Uttar Pradesh might be sharing a low % here but we can attribute it to the low level of awareness of filing a case. I am trying to incoporate these factors but will have to get more data
3) The following graph looks at the age group of the people who commit crimes. As expected most of them lie in the 18-30 range
4) On trying to find out whether rural/urban has a part to play, it was found that the rural graph has more correlation than the urban graph

5)  I regressed the crime occurrences with the following factors
POPULATION  
GROWTH  
RURAL  
URBAN  
AREA 
DENSITY  
SEXRATIO

 The following are the results of the regression. We can see that the R Square is too less to say anything definite about the relationship of crime and the geographical boundaries. There are other factors


ALL INDIA
Age GroupR- Squared
<18 td="">5.82%
18 to 3016.76%
30 to 4514.86%
45 to 6013.97%
> 6011.77%




NORTH STATES
Age Group R- Squared
<18 td=""> 4.55%
18 to 30 7.38%
30 to 45 5.95%
45 to 60 5.70%
> 60 3.34%