Hadoop streaming will run the process in “local” mode when there is no hadoop instance running on the box. I have a shell script that is controlling a set of hadoop streaming jobs in sequence and I need to condition copying files from HDFS to local depending on whether the jobs have been running locally or not. Is there a standard way to accomplish this test? I could do a “ps aux | grep something” but that seems ad-hoc.
Share
Rather than trying to detect at run time which mode the process is operating, it is probably better to wrap the tool you are developing in a bash script that explicitly selects local vs cluster operatide. The O’Reilly Hadoop describes how to explicitly choose local using a configuration file override:
where
conf-local.xmlis an XML file configured for local operation.