Tuesday, August 18, 2015
Increasing the amount of inotify watchers
echo fs.inotify.max_user_watches=524288 | sudo tee -a /etc/sysctl.conf && sudo sysctl -p
Saturday, August 15, 2015
The famous NA12878 WGS at 50x
Due to the lacking of true variants from real genome sequencing, many tools use NA12878 for benchmarking.
For whole genome sequencing at 50x depth, we can download these two files:
ftp.sra.ebi.ac.uk/vol1/fastq/ERR194/ERR194147/ERR194147_1.fastq.gz
ftp.sra.ebi.ac.uk/vol1/fastq/ERR194/ERR194147/ERR194147_2.fastq.gz
Some statistics of these two files| Facts | ERR194147_1.fastq.gz | ERR194147_2.fastq.gz |
|---|---|---|
| URI | ERR194147_1.fastq.gz | ERR194147_2.fastq.gz |
| Size | 48G | 49G |
| Reads# | 787265109 | 787265109 |
| Reads Length | 101 | 101 |
| Coverage | 101*787265109/3000000000=26.5 | 101*787265109/3000000000=26.5 |
Friday, June 12, 2015
Download sequencing data from illumina's basespace
Spent 30 mins creating a script to download sequencing data from illumina's BaseSpace, by using its REST API
Here is the code:
https://gist.github.com/anonymous/91c5accc3988cb233347
Here is the code:
https://gist.github.com/anonymous/91c5accc3988cb233347
Tuesday, May 12, 2015
Simple SAMBA
sudo apt-get install -y samba
mkdir /project/share
chmod -R a+rxw /project/share
###############################
#/etc/samba/smb.conf
[share]
path = /project/share
available = yes
guest only=yes
read only = no
browseable = yes
public = yes
writable = yes
###############################
sudo service smbd restart
\\10.2.5.212\share
mkdir /project/share
chmod -R a+rxw /project/share
###############################
#/etc/samba/smb.conf
[share]
path = /project/share
available = yes
guest only=yes
read only = no
browseable = yes
public = yes
writable = yes
###############################
sudo service smbd restart
\\10.2.5.212\share
Monday, April 20, 2015
Run a BWA job with docker
To run a BWA mappping job inside a docker, I want to create three containers. one for "bwa" executable, one for "bwa" genome index, and one for the input fastq files. All the benefits can be summarized by one word: "isolation".
- 1. The bwa executable application. I would put it into a volume /bioapp/bwa/0.7.9a/ in an docker image named "yings/bioapp"
- 2. The reference genome index which was created using "bwa index". I would put it into a volume /biodata/hg19/index/bwa/ in a docker image named "yings/biodata"
- 3. The input FASTQ files. Assuming they can be found under "/home/hadoop/fastq" in the host.
Create a image name "bioapp" with tag "v04202015"
copy all files and folders under "app" to the "/bioapp/" in the container. For simplicity here other operations like installing dependencies were not included in the Dockerfile.
The Dockerfile looks like this:
FROM ubuntu:14.04
RUN mkdir -p /bioapp
COPY app /bioapp/
VOLUME /bioapp
ENTRYPOINT /usr/bin/tail -f /dev/null
build the bioapp image -$docker build -t yings/bioapp:v04202015 .Create a image name "biodata" with tag "v04202015" copy all files and folders under "data" to the "/biodata/" in the container$cat >Dockerfile< FROM ubuntu:14.04 RUN mkdir -p /biodata COPY data /biodata/ VOLUME /biodata ENTRYPOINT /usr/bin/tail -f /dev/null EOFBuild the bioapp image$docker build -t yings/biodata:v04202015 .Start the biodata container as daemon, name it as "biodata"$docker run -d --name biodata yings/biodata:v04202015Start the bioapp container as daemon, name it as "bioapp"$docker run -d --name bioapp yings/bioapp:v04202015Now we should have two data volume containers running in the backend. It is time to launch the final executor container$docker run -it --volumes-from biodata --volumes-from bioapp -v /home/hadoop/fastq:/fastq ubuntu:14.04 /bin/bashThose parameters mean: "-it" run the executor container interactively "--volumes-from biodata" load the data volume from container "biodata" (do not confused it with image "yings/biodata") "--volumes-from bioapp" load the data volume from container "bioapp" (again, do not confused it with image "yings/bioapp") "-v /home/hadoop/fastq:/fastq ubuntu:14.04" mount the volume "/home/hadoop/fastq" in host to "/fastq" in the executor container. "ubuntu:14.04" this is the standard image, as our OS "/bin/bash" command to be executed as entry point. If everything goes well, you will see you are root now in the executor containerroot@5927eecc8530:/#Is bwa there?root@5927eecc8530:/# ls -R /bioapp/Is genome index there?root@5927eecc8530:/# ls -R /biodata/Is fastq there?root@5927eecc8530:/# ls -R /fastq/Launch the job, save the alignment as "/fastq/output/A.sam"root@5927eecc8530:/#/bioapp/bwa/0.7.9a/bwa mem -t 8 -R '@RG\tID:group_id\tPL:illumina\tSM:sample_id' /biodata/index/hg19/bwa/ucsc.hg19 /fastq/A_R1.fastq.gz /fastq/A_R2.fastq.gz > /fastq/A.samThe process should be complete in a few minutes because the input fastq files are very small, as a test. Now you can safely terminate the executor container by pressing "Ctrl-D". Since we previously mount the volume "/home/hadoop/fastq" in host to "/fastq" in the executor container, now back to the host, we will see the persistent output "A.sam" under "/home/hadoop/fastq". However, if we use the "/biodata" or "/bioapp" as the output folder, for example, "bwa ... > /biodata/A.sam", then it is NOT persistent. If you terminate the "biodata" container, all the changes on that container will be lost. (stop & restart the container is OK, as long as it was not terminated).
Friday, April 17, 2015
Docker container to host reference application and library?
Recently I am considering to migrate from AMI to Docker Image
for the reference repository. Thinking about launching 20 or more on-demand AWS EC2 c3.4xlarge instances for WGS data analysis.
===AMI===
Pros:
Fast to start
No installation required
No configuration required
Secure
Cons:
AWS only
Building AMI is a little time-consuming
Version control is tricky
Volume size will increase gradually
===Docker image===
Pros:
Easy to build
Deploy on any cloud or local platforms
Easy to tag or version control
Seems more popular and cooler
Cons:
Where to host repository? The Hub or S3? Many considerations
Pulling images from one repository on lots of worker nodes at the same time is not applicable. Networking would be challenging.
Definitely need more time to do lots of tests. Work in progress...
for the reference repository. Thinking about launching 20 or more on-demand AWS EC2 c3.4xlarge instances for WGS data analysis.
===AMI===
Pros:
Fast to start
No installation required
No configuration required
Secure
AWS only
Building AMI is a little time-consuming
Version control is tricky
Volume size will increase gradually
===Docker image===
Pros:
Easy to build
Deploy on any cloud or local platforms
Easy to tag or version control
Seems more popular and cooler
Cons:
Where to host repository? The Hub or S3? Many considerations
Pulling images from one repository on lots of worker nodes at the same time is not applicable. Networking would be challenging.
Definitely need more time to do lots of tests. Work in progress...
Friday, April 10, 2015
Lightweight MRAppMaster?
hadoop v2 moves the master application from Master Node to one container. The pro is that it greatly reduced the worload on Master Node while launching multiple MapReduce applications and easily make the hadoop framework more scalable beyond thousands of Worker Nodes.
However this brings a problem in a small cluster- the Master Application "MRAppMaster" now will occupy a whole container.
Supposedly we have five Worker Nodes and every node has just one
container. Since "MRAppMaster" will use one container, then you have only four containers can be used for real processing. 20% of computing resource were "wasted". We can mitigate this problem by assigning two containers per node. By this way only 10% of computing resource were wasted. However if we divide a node into two containers, then every container's memory and CPU will be cut into half too. The memory size is very precious in many bioinformatics applications. With less than 8G memory your aligner probably will fail.
If your hadoop cluster has more than 10 nodes, then do not bother to take it into consideration.
However this brings a problem in a small cluster- the Master Application "MRAppMaster" now will occupy a whole container.
Supposedly we have five Worker Nodes and every node has just one
container. Since "MRAppMaster" will use one container, then you have only four containers can be used for real processing. 20% of computing resource were "wasted". We can mitigate this problem by assigning two containers per node. By this way only 10% of computing resource were wasted. However if we divide a node into two containers, then every container's memory and CPU will be cut into half too. The memory size is very precious in many bioinformatics applications. With less than 8G memory your aligner probably will fail.
If your hadoop cluster has more than 10 nodes, then do not bother to take it into consideration.
Subscribe to:
Posts (Atom)